Skip to content

Batch Guardian log writes and drain them on shutdown - #112

Open
rksharma-owg wants to merge 1 commit into
GenAI-Security-Project:integrationfrom
rksharma-owg:refimpl/90-span-batching
Open

Batch Guardian log writes and drain them on shutdown#112
rksharma-owg wants to merge 1 commit into
GenAI-Security-Project:integrationfrom
rksharma-owg:refimpl/90-span-batching

Conversation

@rksharma-owg

Copy link
Copy Markdown

What changed

Guardian envelope and session-context logs currently call appendFileSync on the request path. Batch their existing JSONL records through bounded asynchronous append streams, and drain both streams after active requests finish during shutdown.

Each sink flushes after 64 records or 100 ms from its first queued record. Both queued and in-flight data count toward limits of 4 MiB and 4,096 records. A failed or full sink reports once and disables itself without changing the policy decision. SIGINT and SIGTERM use the same drain path as guardian.close().

Log schemas and record ordering stay the same, but log visibility is delayed. Abrupt termination can lose buffered records, and graceful shutdown has no forced deadline. The Guardian README documents these limits. This implements batching for existing reference-implementation logs; OpenTelemetry collection/export remains separate work in #91.

Which issue does this implement

Closes #90

Base branch

  • integration, because this touches a reference implementation and tests

Type of change

  • Reference implementation or adapter
  • Documentation

I tested this

  • I synced my branch with the base branch before opening this
  • uv run pytest -v: 278 passed, 1 skipped
  • uv run mkdocs build --strict: passed
  • From reference-implementations/agt, bun run typecheck: passed
  • From reference-implementations/agt, bun test with Bun 1.3.11: 1,122 passed, 1 skipped, 0 failed

The burst-buffering regression fails against the synchronous implementation and passes with this change. Tests cover batch thresholds, sparse traffic, record order, UTF-8 byte and record limits including in-flight writes, write failures, duplicate error reporting, active-request shutdown, and standalone SIGINT/SIGTERM shutdown.

The existing Bun skip requires UPSTREAM_BUNDLE for a pinned Rego byte-identity comparison. The Python skip requires a case-sensitive filesystem. Both also skipped on the baseline.

Checklist

  • Commits are signed off with git commit -s (required by the DCO)
  • Prose follows STYLE.md
  • No secrets, tokens, or internal URLs in the diff

Security

This changes log buffering and shutdown durability. It does not change policy decisions or report a vulnerability. Abrupt termination and sink errors can lose buffered log records, as documented; buffering is bounded to limit memory use.

Signed-off-by: RKS <rajesh.sharma@owasp.org>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add span batching to the reference implementation

1 participant