Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart · Issue #13578 · objectstack-ai/objectstack · GitHub
Skip to content

datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578

Description

@os-steve

Split out of #13408 by the domain:cli PM seat (#6024) at dispatch time, unlabelled for triage. Filed separately because it is decision-independent and #13408 is not, see below.

The defect

On a multi-node deployment, a datasource whose driver fails to start leaves a stuck driver instance in the in-memory data-engine driver registry. Deleting the datasource does not clear it:

  • DELETE /api/v1/datasources/:name → the admin-door list is empty on every replica…
  • …but GET /api/v1/readystill names that datasource's driver and still answers 503.
  • No API door evicts the engine-registry entry. Only a process restart clears it.

So an administrator who reacts to a bad datasource with the single most natural action — delete it — gets no recovery, and on a shared/HA cluster the remaining remedy is a restart of every replica.

Why this is filed apart from #13408, and not held behind it

#13408 bundles two halves. Its other half — should one unhealthy driver drain whole-node readiness? — turned out to be intentional, documented behaviour, so it is a product decision and #13408 has been moved to the decision box (analysis on that card).

This half is not that. It is a defect under either answer to the readiness question:

  • If a bad driver should drain the node, then a deleted datasource must stop draining it — otherwise the drain is unrecoverable by design.
  • If it should not, this is still a registry that leaks entries no door can reach.

⇒ Holding it behind the decision would park a live defect inside a decision it does not depend on. Recording that reasoning here so the split is auditable rather than looking like scope drift.

What the fix needs to answer

The repair is not only the DELETE path — that is the one path that happened to be observed:

  • Enumerate every path that can leave an orphan driver instance, by walking the registry's lifecycle rather than by example: failed-start rollback, tenant deletion, datasource rename/reconfigure, environment teardown. ⛔ An answer of the form "fixed DELETE" leaves the others.
  • Decide where eviction belongs — the delete path calling the registry, or the registry owning its own liveness — and say why.
  • ⚠️Cluster propagation is part of the shape, not a detail.[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records the sibling asymmetry on a different registry: create broadcasts cluster-wide, delete does not mirror. If eviction is local-only, one replica recovers and the rest do not, which is the same bug with a smaller blast radius.

Landing surface

The data-engine driver registry plus the datasource delete path — notpackages/runtime/src/http-dispatcher.ts, which only reports the registry's contents at /ready. #13408's triage note flagged this half as needing engine/services alignment; routing is triage's call, which is why this card is filed unlabelled.

Refs

Observed on a live 3-replica EE deployment during the #13404 checklist run; recovered by restart.

Metadata

Metadata

Assignees

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions