fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix: atexit final flush + dispose-during-connect (closes #6) - #28

Merged
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle
May 13, 2026
Merged

fix: atexit final flush + dispose-during-connect (closes #6)#28
AndresL230 merged 2 commits into
mainfrom
fix/process-lifecycle

Conversation

@AndresL230

@AndresL230AndresL230 commented May 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes#6.

Two surgical fixes to process lifecycle:

  • atexit final flushinit() registers handle.dispose() via
    atexit so short-lived processes (cron, Lambda, one-shot CLI, SIGTERM
    in a container) flush their last bucket instead of dropping it when
    the daemon flush thread is killed. Opt-out via
    RecostConfig.auto_shutdown_handlers=False.
  • dispose during connect_LocalTransport.dispose() now schedules
    loop.stop() via call_soon_threadsafe instead of waiting for a
    queue sentinel. If the loop is blocked inside websockets.connect()'s
    upgrade handshake, the sentinel-in-queue dispose used to leave the
    daemon thread + socket FD pinned until the OS TCP timeout (~75 s).

Tests

  • + "" + tests/test_transport.py::TestDisposeDuringConnect + "" + — stands up a
    TCP server that accepts the connection but never responds to the
    WebSocket upgrade, then calls + "" + dispose() + "" + mid-handshake. Post-fix
    dispose returns in well under one second (threshold 1.5 s).
  • + "" + tests/test_init.py::TestAtexitFlush + "" + — three tests: atexit IS
    registered when enabled, atexit is NOT registered when opted out,
    and an end-to-end subprocess test that verifies a real + "" + sys.exit(0) + "" +
    triggers the final flush.

Notes

  • Public API unchanged. New + "" + RecostConfig.auto_shutdown_handlers + "" +
    defaults to + "" + True + "" + — existing callers automatically get the better
    behavior with no migration.
  • The atexit callback unregisters itself on explicit + "" + dispose() + "" + so
    long-lived processes that cycle + "" + init + "" + / + "" + dispose + "" + don't accumulate
    dead callbacks.
  • Builds on PR fix: guard module-level state with RLock (closes #4) #27 (issue Module-level state races: _handle, install/uninstall, init-vs-dispose #4) — + "" + _init_lock + "" + already serializes
    + "" + init + "" + / + "" + dispose + "" + , so atexit re-entering + "" + dispose() + "" + from the main
    thread is safe against user-driven dispose.
  • Python 3.14 compatibility: + "" + loop.stop() + "" + on the loop thread now
    causes + "" + run_until_complete + "" + to raise
    + "" + RuntimeError: Event loop stopped before Future completed + "" +
    caught with a narrow message match so genuine coroutine bugs still
    surface.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic shutdown handlers to ensure telemetry is flushed when the process exits (configurable and enabled by default).
  • Bug Fixes

    • Resolved thread and socket leak issues during shutdown.
    • Added configurable shutdown timeout to prevent hangs during graceful termination.

Review Change Stack

The old _LocalTransport.dispose() queued a None sentinel and joined with
a 2s timeout. If the loop was inside websockets.connect()'s blocking
upgrade handshake, the sentinel sat in the queue until the OS TCP
timeout (~75 s on Linux), so the join timed out and the daemon thread
plus its open socket FD leaked.
Schedule loop.stop() via call_soon_threadsafe instead — this interrupts
the connect coroutine and lets run_until_complete return promptly.
Bound the join with shutdown_flush_timeout_ms (wired through Transport)
and warn (not block) if the thread is still alive after that.
Regression test stands up a TCP server that accepts but never completes
the WebSocket upgrade, then dispose()s mid-handshake and asserts the
join returns in under 3 s.
Refs #6
Short-lived processes (cron jobs, Lambda invocations, one-shot CLI
scripts, SIGTERM'd containers) dropped their last aggregator bucket
because the flush timer is a daemon thread and dies on exit. The user
either had to remember to call handle.dispose() manually or accept
silent data loss for the final window.
Register handle.dispose() via atexit when init() runs, gated by the
new RecostConfig.auto_shutdown_handlers (default True). The callback
is idempotent and unregisters itself on explicit dispose so a process
that cycles init/dispose does not accumulate dead atexit hooks.
Regression tests: two in-process checks (registered when enabled,
NOT registered when opted out) plus a subprocess end-to-end test that
verifies a normal sys.exit(0) triggers the final flush.
Refs #6
@coderabbitai

coderabbitaiBot commented May 13, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR addresses process lifecycle gaps by implementing graceful shutdown through atexit registration and fixing resource leaks during dispose. It adds a configurable shutdown timeout to prevent FD accumulation in long-lived processes, registers automatic final flushes on normal exit for short-lived processes, and provides comprehensive test coverage including subprocess and blocking-connect scenarios.

Changes

Process Lifecycle and Transport Shutdown

Layer / File(s)Summary
Configuration flag for auto-shutdown behavior
recost/_types.py
RecostConfig.auto_shutdown_handlers (default True) controls whether init() registers an atexit handler for final flush on process exit.
Transport shutdown with configurable timeout
recost/_transport.py
_LocalTransport accepts shutdown_timeout_s parameter and uses it to bound thread.join(). Shutdown via loop.stop() (instead of queue sentinel) handles asyncio.CancelledError cleanly and prevents hangs during websocket handshake. Transport.__init__ wires config.shutdown_flush_timeout_ms / 1000.0 into the transport.
Atexit handler registration and cleanup
recost/_init.py
RecostHandle tracks the registered atexit callback in _atexit_callback field. init() registers handle.dispose via atexit.register() when enabled; dispose() unregisters and clears the callback to keep teardown idempotent across repeated init/dispose cycles.
Atexit and dispose-during-connect test coverage
tests/test_init.py, tests/test_transport.py
TestAtexitFlush suite verifies init() registers callbacks when enabled and skips when disabled. End-to-end subprocess test injects metrics with long flush interval, exits normally, and asserts the final atexit flush writes a marker. TestDisposeDuringConnect verifies _LocalTransport.dispose() returns promptly when called during blocking websockets.connect() handshake.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 Processes now gracefully exit with a final flush,
No more daemon threads in a leaked-FD rush,
Timeouts bound the shutdown, callbacks all cleaned,
Short scripts and long services finally serene!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 57.14% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'fix: atexit final flush + dispose-during-connect (closes #6)' is highly specific and accurately describes the two main fixes addressed in the changeset: implementing atexit-based final flush and resolving the dispose-during-connect thread/FD leak issue.
Linked Issues check✅ PassedAll primary coding objectives from issue #6 are met: atexit registration with auto_shutdown_handlers flag [#6], bounded dispose-during-connect handling via loop.stop() and timeout [#6], and comprehensive test coverage verifying both features and regression scenarios.
Out of Scope Changes check✅ PassedAll changes directly address issue #6 requirements: atexit callback infrastructure, transport shutdown improvements, configuration flag addition, and targeted regression tests—no unrelated modifications detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/process-lifecycle

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/test_transport.py`:
- Around line 417-437: The test helper _start_blackhole_server currently
accumulates accepted client sockets in the local accepted list and never closes
them, leaking file descriptors; modify the implementation so accepted sockets
are closed during teardown by either (a) attaching the accepted list to the
returned srv object (e.g. srv._accepted = accepted) or returning a tuple (srv,
accepted) and then updating tests to iterate over accepted and call .close()
before/after closing srv, and also ensure accept_loop closes any connections
when detecting srv.fileno() == -1 (or on shutdown) to avoid leaving sockets
open; update references to accept_loop, accepted, and _start_blackhole_server
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 574fcaa6-ac40-41fe-b206-7f9771016e62

📥 Commits

Reviewing files that changed from the base of the PR and between 6992c21 and 10a8e51.

📒 Files selected for processing (5)
  • recost/_init.py
  • recost/_transport.py
  • recost/_types.py
  • tests/test_init.py
  • tests/test_transport.py

Comment on lines +417 to +437
def _start_blackhole_server(port: int) -> socket.socket:
"""Bind+listen on `port` and accept connections but never respond
to the HTTP upgrade. websockets.connect() will TCP-connect, send its
upgrade request, and then block waiting for an HTTP response."""
srv = socket.socket()
srv.bind(("127.0.0.1", port))
srv.listen(8)
accepted: list[socket.socket] = []

def accept_loop() -> None:
srv.settimeout(0.5)
while True:
try:
conn, _ = srv.accept()
accepted.append(conn)
except (OSError, socket.timeout):
if srv.fileno() == -1:
return

threading.Thread(target=accept_loop, daemon=True).start()
return srv

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Close accepted blackhole sockets during teardown.

At Line 431, accepted client sockets are retained but never closed; Line 462 only closes the listening socket. This can leak FDs across tests.

Proposed fix
- def _start_blackhole_server(port: int) -> socket.socket:+ def _start_blackhole_server(port: int) -> tuple[socket.socket, list[socket.socket]]:
@@
- return srv+ return srv, accepted
@@
- server = self._start_blackhole_server(port)+ server, accepted = self._start_blackhole_server(port)
@@
finally:
+ for conn in accepted:+ try:+ conn.close()+ except OSError:+ pass
server.close()

Also applies to: 461-463

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_transport.py` around lines 417 - 437, The test helper
_start_blackhole_server currently accumulates accepted client sockets in the
local accepted list and never closes them, leaking file descriptors; modify the
implementation so accepted sockets are closed during teardown by either (a)
attaching the accepted list to the returned srv object (e.g. srv._accepted =
accepted) or returning a tuple (srv, accepted) and then updating tests to
iterate over accepted and call .close() before/after closing srv, and also
ensure accept_loop closes any connections when detecting srv.fileno() == -1 (or
on shutdown) to avoid leaving sockets open; update references to accept_loop,
accepted, and _start_blackhole_server accordingly.

@AndresL230
AndresL230 merged commit edf8b23 into mainMay 13, 2026
1 check passed
@AndresL230
AndresL230 deleted the fix/process-lifecycle branch May 21, 2026 04:17
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process lifecycle: no atexit flush, no signal handlers, dispose can leak threads

1 participant

@AndresL230