feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(observability): admin analytics + cost API /api/admin/analytics (#120) - #376

Merged
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api
Jul 27, 2026
Merged

feat(observability): admin analytics + cost API /api/admin/analytics (#120)#376
Darkest-Teddy merged 1 commit into
feat/118-llm-usage-observabilityfrom
feat/120-admin-analytics-api

Conversation

@Darkest-Teddy

Copy link
Copy Markdown
Collaborator

What & why

Closes#120 — the admin-only, read-only analytics + cost-rollup API over the events / llm_usage tables. This is the backend the future admin dashboard (#121/#122) consumes.

New routes/admin_analytics.py, mounted at /api/admin/analytics, every endpoint gated by require_admin.

Endpoints

MethodPathReturns
GET/usage/summaryper-event_type counts, total events, distinct active users
GET/usage/by-userper-user event counts (by category) + LLM cost_usd + total_tokens, paginated
GET/llm/cost?group_by=user|feature|modeltoken + cost rollups by dimension, with totals
GET/errorserror.* events with path/method/status/duration from payload, paginated

All accept from/to ISO bounds (default: last 30 days) and return typed Pydantic models. No raw content — fingerprints only.

Notes

  • PostgREST has no GROUP BY, so the grouped endpoints scan the date-bounded rows and aggregate in Python (as the issue sanctions). Scans page via select_with_count and use its exact count both to stop and to detect the _SCAN_CAP truncation ceiling — logged, never a silent row-cap loss. /errors needs no aggregation, so it paginates server-side.
  • Date-range filtering uses a list-valued created_at param (gte.… + lte.…), which PostgREST ANDs.

Success criteria (#120)

  • All four endpoints return correct aggregates against seeded rows (tested).
  • Every endpoint rejected for non-admins / allowed for admins (require_admin, tested).
  • group_by switches user/feature/model correctly.
  • Date-range filtering + pagination work; default range last 30 days.
  • Router mounted in main.py under /api/admin/analytics.
  • Tests in backend/tests/, pass with a faithful mocked table().

Testing

11 cases over a fake table() that honors eq/gte/lte/like filters, ordering, and limit/offset — so aggregation, group_by switching, date filtering, default range, pagination, and admin gating are all genuinely exercised.

Stacking

Stacked on #375 (#118) — needs the 0032events/llm_usage tables. Base branch is feat/118-llm-usage-observability; retarget to main once #375 merges.

🤖 Generated with Claude Code

…120)
Closes#120. Read-only, admin-only query API over the events + llm_usage
tables (from #118) — the backend the future admin dashboard consumes.
New routes/admin_analytics.py, mounted at /api/admin/analytics, every
endpoint gated by require_admin:
- GET /usage/summary — per-event_type counts, total events, distinct
active users in range.
- GET /usage/by-user — per-user event counts (by category) + LLM
cost_usd + total_tokens; paginated.
- GET /llm/cost?group_by= — token + cost rollups grouped by user|feature|
model, with totals.
- GET /errors — error.* events with path/method/status/duration
from payload; paginated server-side.
All endpoints take from/to ISO bounds (default last 30 days) and return
typed Pydantic models (fingerprints only, no raw content). PostgREST has no
GROUP BY, so grouped endpoints scan the date-bounded rows and aggregate in
Python; scans page via select_with_count and use its exact count to stop and
to detect (and log) the truncation ceiling — no silent row-cap loss.
Tests: 11 cases over a faithful table() fake covering aggregation, group_by
switching, date-range filtering, default range, pagination, and admin gating.
Stacked on #118 (needs the 0032 events/llm_usage tables).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^production$
  • ^staging$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: c720c9b7-9665-4141-8916-1b657f906f0d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login.

Comment @coderabbitai help to get the list of available commands.

@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 22, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-staging7f1e776Commit Preview URL

Branch Preview URL
Jul 22 2026, 12:52 AM

@Darkest-Teddy
Darkest-Teddy merged commit be376eb into feat/118-llm-usage-observabilityJul 27, 2026
4 checks passed
AndresL230 added a commit that referenced this pull request Jul 29, 2026
… admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 29, 2026
* feat(observability): capture LLM token usage + cost per call (#118) + admin analytics (#120)
Rebase of PR #375 (which included stacked #376) onto current main, rebuilt as
a fresh application of the branch diff plus the adaptations main now requires.
Core (unchanged from #375/#376):
- agents/usage.py: record_agent_usage(result, feature=, task=, user_id=) —
one-line, never-raising usage capture for every Pydantic AI run.
- services/events_service.py: bounded-queue fire-and-forget writer draining
events/llm_usage rows through db/connection.table() off the request thread.
- services/llm_pricing.py: token-field normalization + per-1K price map;
unknown real models record cost_usd=NULL with a one-time warning.
- gemini_service: call_gemini / call_gemini_multiturn log via
_log_gemini_usage (feature= threaded from callers; json delegates).
- routes/admin_analytics.py (+ mount): admin usage/cost rollup endpoints.
- tests/test_usage_instrumentation_coverage.py: AST guard — a module that
runs an agent without referencing record_agent_usage fails CI.
Rebase adaptations:
- Migration renumbered 0032_observability.sql -> 0035_observability.sql
(main grew 0032-0034 in the meantime); content unchanged, header comment
updated.
- main.py lifespan: events_service.start_worker()/shutdown() coexists with
#406's Logfire wiring (configure/instrument_pydantic_ai/instrument_fastapi).
- Conflict resolutions keep main's guardrail semantics and add capture on
top: notes.py wraps _run_note_worker results (WORKER_LIMITS + 413/500
mapping intact) and the note_chat try/except (ORCHESTRATOR_LIMITS +
degraded reply intact); learn.py wraps the _prepare_chat_run-based
_chat_via_agent; calendar_service keeps usage_limits=WORKER_LIMITS.
- NEW run-sites landed on main since the branch:
* SSE streaming tutor (#349): stream_agent_turn grows an optional
on_usage(run_result) hook fed by the final AgentRunResultEvent, called
once on the success path before on_complete (tokens are spent even if
persistence fails); /chat/stream and /start-session/stream pass
record_agent_usage(feature="chat_tutor", task="chat_tutor"). Error rungs
and the Rung-1 legacy fallback don't fire it — legacy usage is captured
inside call_gemini_multiturn(feature=). Covered in test_chat_stream.py.
* gemini_vision_backend: per-page ocr_vision_agent run wrapped
(feature="document", task="ocr_vision"; no user_id — extraction is
content-addressed and user-agnostic).
- SAPLING_MODEL_MODE=function: 'function:<task>' model names record with
cost_usd=NULL and NO unpriced-model warning (they are the e2e/CI seam,
not real spend); real unknown models keep the one-time warning. Tested.
Verification: full backend suite 1275 passed / 27 skipped; ruff clean; AST
guard green over all current run-sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(observability): apply #375 review fixes — range validation, truncation flag, private caching, poison-row salvage, per-call-site usage guard
- admin_analytics _resolve_range: validate from/to as ISO 8601 (422 naming
the bad param) and reject from > to; validated strings echoed unchanged.
- Surface the 100k _SCAN_CAP: _scan_range returns (rows, truncated) and
UsageSummary/UsageByUser/LLMCost carry `truncated: bool = False` (the
/errors feed paginates server-side, so it has no cap to surface).
- All 4 analytics GETs now send `Cache-Control: private` via a Response param.
- test_default_range_is_last_30_days freezes the module clock at 2026-07-21
so the fixture window can't rot after 2026-08-09.
- events_service._flush_batch: a failed bulk insert now retries rows one at
a time, dropping only the rows that individually fail (per-row debug log +
one warning with the drop count); never-raise contract kept, unit-tested
with a 3-row batch where only the poison row is lost.
- test_usage_instrumentation_coverage: upgraded the file-level substring
guard to a real per-call-site AST check — every agent-run site needs an
enclosing record_agent_usage, with pass-through runner helpers (return
await agent.run(...)) checked at their module-local callers instead. The
docstring now states the exact granularity and the cross-module blind spot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: AndresL230 <190146319+AndresL230@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/120-admin-analytics-api branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@Darkest-Teddy