Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Fix/add connection pools to fix db hangup - #731

Merged
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup
Jan 27, 2026
Merged

Fix/add connection pools to fix db hangup#731
coodos merged 3 commits into
mainfrom
fix/add-connection-pools-to-fix-db-hangup

Conversation

@coodos

@coodoscoodos commented Jan 27, 2026

Copy link
Copy Markdown
Contributor

Description of change

Issue Number

closes#638

Type of change

  • Breaking (any change that would cause existing functionality to not work as expected)
  • New (a change which implements a new feature)
  • Update (a change which updates existing functionality)
  • Fix (a change which fixes an issue)
  • Docs (changes to the documentation)
  • Chore (refactoring, build scripts or anything else that isn't user-facing)

How the change has been tested

Change checklist

  • I have ensured that the CI Checks pass locally
  • I have removed any unnecessary logic
  • My code is well documented
  • I have signed my commits
  • My code follows the pattern of the application
  • I have self reviewed my code

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Improved webhook participant loading to avoid hangs and timeouts, with safer per-item handling and better failure isolation.
    • Reworked change handling to reliably reload and send updated entities after commits, reducing missed or inconsistent notifications.
  • Performance

    • Tuned database connection pooling and timeouts to improve resource utilization and responsiveness.

✏️ Tip: You can customize this high-level summary in your review settings.

The subscriber's afterUpdate was using event.entity which only contains
changed fields (partial entity), not the full entity with charter. When
groups were updated via webhooks, the charter field was often absent in
the partial entity, causing Cerberus to lose track of group charters
over time. After restart, charters would be loaded fresh from DB.
Root cause: TypeORM's afterUpdate event provides partial entities when
using repository.save(entity). The code was not reloading the full
entity with all fields (including charter) from the database after the
transaction committed.
Changes:
- Refactored afterUpdate to pass metadata (entityId, relations) instead
of using the partial entity from event.entity
- Created handleChangeWithReload that schedules entity reload for after
transaction commit (inside setTimeout with 50ms delay)
- Created executeReloadAndSend that does the actual findOne with all
relations AFTER transaction commits, ensuring charter and all fields
are loaded
- Groups and messages sync with 50ms delay (fast, ensures commit)
This is the same transaction timing issue fixed in file-manager-api and
dreamsync-api. Now Cerberus will maintain full group data (including
charters) indefinitely without requiring restarts.
The group webhook processing could hang indefinitely when loading
participant users, causing Cerberus to appear stuck after running
for some time. The last log was "Extracted userId" with no progress.
Root causes:
1. Promise.all blocks if any getUserById call hangs (DB lock, timeout, etc.)
2. No timeout protection - hangs wait forever
3. No error handling - failures block entire webhook
4. Loading unnecessary relations (followers/following) added complexity
Changes:
- Use Promise.allSettled instead of Promise.all to handle failures gracefully
- Add 5-second timeout per user lookup using Promise.race
- Wrap each participant load in try-catch with detailed error logging
- Load users without heavy relations in webhook context (don't need followers/following)
- Add indexed logging to identify which participant causes issues
- Log success/failure counts for transparency
Benefits:
- Webhook completes even if some participants fail to load
- 5s timeout prevents indefinite hangs
- Better diagnostics via indexed logging
- Reduced DB load by skipping unnecessary relations
The webhook will now respond within ~5 seconds even if all participant
lookups fail, preventing Cerberus from getting stuck.
@coderabbitai

coderabbitaiBot commented Jan 27, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Enhances webhook participant loading with per-participant timeouts to avoid hangs, adds connection-pool tuning to multiple Postgres DataSource configs, and refactors the subscriber afterUpdate flow to reload, enrich, and debounce post‑commit entity notifications.

Changes

Cohort / File(s)Summary
Webhook participant loading
platforms/cerberus/src/controllers/WebhookController.ts
Replaced prior per-participant async loading with a per-item async function that safely extracts userId, wraps each repository call in a 5s Promise.race timeout, and uses Promise.allSettled to collect valid participants; adds per-index error logging.
Subscriber reload-and-notify workflow
platforms/cerberus/src/web3adapter/watchers/subscriber.ts
Reworked afterUpdate handling to derive entityId from multiple sources and schedule a debounced post‑commit reload. Added private handleChangeWithReload and executeReloadAndSend methods that reload/enrich the entity, convert to plain data, skip junctions/non-system messages, and call adapter.handleChange. Includes additional guards and logging.
Database connection pool configs
platforms/cerberus/src/database/data-source.ts, infrastructure/evault-core/src/config/database.ts, platforms/dreamsync-api/src/database/data-source.ts, platforms/eCurrency-api/src/database/data-source.ts, platforms/eReputation-api/src/database/data-source.ts, platforms/emover-api/src/database/data-source.ts, platforms/esigner-api/src/database/data-source.ts, platforms/evoting-api/src/database/data-source.ts, platforms/file-manager-api/src/database/data-source.ts, platforms/group-charter-manager-api/src/database/data-source.ts, platforms/pictique-api/src/database/data-source.ts, platforms/registry/src/config/database.ts
Added an extra block to DataSource/DataSource options across multiple platforms with connection-pool and timeout settings (max: 10, min: 2, idleTimeoutMillis: 30000, connectionTimeoutMillis: 5000, statement_timeout: 10000). Review for consistent naming/typing and environment compatibility.

Sequence Diagram(s)

sequenceDiagram
participant Event as Change Event
participant Sub as Subscriber
participant DB as Database
participant Adapt as Adapter
Event->>Sub: afterUpdate(event)
Sub->>Sub: derive entityId (event.entity.id / databaseEntity?.id / common id fields)
alt id missing
Sub-->>Event: log warning & exit
else id present
Sub->>Sub: schedule handleChangeWithReload (debounced per table)
Note right of Sub: debounced timer per table (skip junctions)
Sub->>DB: reload entity by id after commit
Note over Sub,DB: reload uses repository.findOne and enrichment
DB-->>Sub: enriched entity/plain data
Sub->>Sub: validate (skip locked/non-system where applicable)
Sub->>Adapt: adapter.handleChange(envelope)
Adapt-->>Sub: acknowledgement
Sub->>Sub: log envelope
end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Bekiboo
  • sosweetham

Poem

🐰
I hopped through logs and timeouts bright,
Per‑participant naps cut short at night,
I fetch, I reload, I debounce with care,
Group chats chirp now — no hangs in the air! ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 2
❌ Failed checks (2 inconclusive)
Check nameStatusExplanationResolution
Linked Issues check❓ InconclusiveThe PR addresses issue #638 (Cerberus group chat failures) through connection pool additions and webhook/subscriber improvements, but the connection between specific code changes and the expected behavior is not explicitly documented.Clarify in the PR description how the connection pool and webhook timeout changes specifically address the notification and charter violation processing failures described in issue #638.
Out of Scope Changes check❓ InconclusiveThe PR contains mostly in-scope changes (database pool configurations and webhook/subscriber enhancements) related to fixing database hangups, but includes additional changes to subscriber reload logic that may extend beyond the immediate scope of connection pooling.Review whether the subscriber reload-and-notify workflow (155 lines added) is necessary for the core hangup fix or if it should be separated into a distinct PR for better clarity.
✅ Passed checks (3 passed)
Check nameStatusExplanation
Title check✅ PassedThe title 'Fix/add connection pools to fix db hangup' directly relates to the main changes in the PR, which add connection pool configurations across multiple database data sources to address database hangup issues.
Description check✅ PassedThe PR description follows the provided template structure with Issue Number (#638) and lists all change type options, but the 'Type of change' section is not properly checked and 'How the change has been tested' is empty.
Docstring Coverage✅ PassedNo functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coodos
coodosforce-pushed the fix/add-connection-pools-to-fix-db-hangup branch from 83dce76 to c05d7d0CompareJanuary 27, 2026 20:59
@coodos
coodos merged commit 46b062e into mainJan 27, 2026
7 checks passed
@coodos
coodos deleted the fix/add-connection-pools-to-fix-db-hangup branch January 27, 2026 21:13
@coderabbitaicoderabbitaiBot mentioned this pull request Mar 16, 2026
6 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Cerberus doesn't work with group chats and charters

2 participants

@coodos@ananyayaya129