Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
 Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Follow-up to #792 (multi-LSP support) by Camillarhi · Pull Request #956 · lightningdevkit/ldk-node · GitHub
Skip to content

Follow-up to #792 (multi-LSP support) - #956

Merged
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up
Jul 21, 2026
Merged

Follow-up to #792 (multi-LSP support)#956
tnull merged 2 commits into
lightningdevkit:mainfrom
Camillarhi:multi-lsp-support-follow-up

Conversation

@Camillarhi

@CamillarhiCamillarhi commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Some minor follow-ups after #792 landed.

  • Retry LSP protocol discovery when it fails at startup, so a transient connection failure doesn't leave a configured LSP permanently unusable. Adds a background retry with backoff and a Liquidity::retry_discovery(node_id) method for on-demand re-discovery (also useful when an LSP rolls out a new protocol).

  • Honor trust_peer_0conf independent of the LSP's supported protocols, It was previously only applied to LSPS2 peers, so LSPS1-only or undiscovered LSPs silently lost their configured 0-conf trust.

Leaving the LSP feature-gating for #900

Fixes - #936

@ldk-reviews-bot

ldk-reviews-bot commented Jul 2, 2026

Copy link
Copy Markdown

👋 Thanks for assigning @tnull as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excuse the delay here. Second commit seems broken, please avoid the unrelated changes as otherwise it's not really reviewable. Also needs a minor rebase by now.

Comment threadsrc/lib.rs Outdated
let retry_ls = Arc::clone(&self.liquidity_source);
let retry_logger = Arc::clone(&self.logger);
let retry_cm = Arc::clone(&self.connection_manager);
self.runtime.spawn_background_task(async move {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That likely should be a cancellable task?

Comment threadsrc/liquidity/mod.rs Outdated
/// The `node_id` must belong to an LSP configured at build time or added via
/// [`Liquidity::add_liquidity_source`]; otherwise [`Error::LiquiditySourceUnavailable`]
/// is returned.
pub fn retry_discovery(&self, node_id: PublicKey) -> Result<(), Error> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a fan of exposing this publicly. Retrying should not be the concern of a user?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was meant to let a user trigger re-discovery for an LSP they know has rolled out a new protocol, the background retry only covers LSPs that never completed discovery in the first place, so an already-discovered LSP adding a protocol wouldn't get picked up.
But thinking about it more, the user usually won't know when an LSP adds a protocol, so leaning on them to call this isn't great. Probably makes more sense to have the background task periodically re-run discovery for already-discovered LSPs too, so new protocols get picked up automatically and there's nothing to expose. I will extend the background check

Comment threadsrc/liquidity/mod.rs Outdated
.collect()
}

pub(crate) fn get_single_lsp_details(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please avoid such ~single-call helpers (especially if they are only used in one place). Please rather inline them so they don't clutter up the file as much.

Comment threadsrc/lib.rs Outdated
loop {
let undiscovered_lsps = retry_ls.get_undiscovered_lsps();
if undiscovered_lsps.is_empty() {
tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm? If there's nothing to do, why do need this select?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is there to stop the loop from busy-spinning when there's nothing to discover. The idea was to keep the task alive in case an LSP became undiscovered later, but a runtime add_liquidity_source cleans up the node if its discovery fails, so nothing new ever lands in the undiscovered set after startup. So once it's empty, there's nothing to wait for, and it should just return. I'll clean it up as part of the re-discovery rework, for re-discovering already-discovered LSPs.

Comment threadsrc/lib.rs Outdated
}
}

tokio::select! {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems an tokio::time::interval would be more suitable?

Comment threadsrc/lib.rs Outdated

for (node_id, address) in undiscovered_lsps {
if let Err(e) =
retry_cm.connect_peer_if_necessary(node_id, address.clone()).await

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this seems redundant to our general reconnection loop, or not?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Configured LSPs aren't added to the peer store at startup, so the reconnection loop doesn't keep them connected. This is what connects us to run discovery.

Comment threadsrc/liquidity/client/mod.rs
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from da5bcdc to b6da9a3CompareJuly 9, 2026 20:38
@Camillarhi
Camillarhi requested a review from tnullJuly 9, 2026 20:42
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from b6da9a3 to 27082aaCompareJuly 10, 2026 12:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two minor comments.

Comment threadsrc/lib.rs
connect_and_discover_lsp(&cm, &ls, &logger, node_id, address).await;
});
}
discovery_set.join_all().await;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Make rediscovery cancellation-safe — /home/tnull/worktrees/ldk-node/pr-956/src/lib.rs:798

    Node::stop() aborts this task while join_all() may be awaiting an LSPS0 request. Dropping discover_lsp_protocols bypasses its timeout cleanup, leaving the node ID permanently present in pending_lsps0_discovery. After restarting the same Node, every discovery attempt reports “already in
    flight,” preventing protocol refresh unless the old response eventually arrives. Use drop-based cleanup for pending requests or allow in-flight discovery tasks to finish during shutdown.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed with drop-based cleanup, a guard now removes the pending entry when the discovery future is dropped, so an aborted task can't leave a stale "in flight" entry behind.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection. And wouldn't peers that never finish discovery also be retried above?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we positive we need this? I think when an LSP updates it's fine to only discover that on next reconnection.

Unused LSPs aren't persisted to the peer store, so the reconnection loop won't cover them, and discovery doesn't currently run on reconnection, but I could watch for when an LSP disconnects and re-run discovery once it reconnects.

And wouldn't peers that never finish discovery also be retried above?

The retry loop above this only covers never-discovered peers until the backoff caps out, after that the sweep is what keeps retrying them. If I swap the sweep for the rediscover-on-reconnect approach, I'd let the retry loop keep going at the cap instead, so those stay covered.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is correct, a restart will cover this. I'll drop the rediscovery for already-discovered LSPs and keep only the retry for the ones where the initial discovery failed, running at the capped backoff until they come up, and stopping once there's nothing left undiscovered

@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 27082aa to 78c313fCompareJuly 15, 2026 13:59
@Camillarhi
Camillarhi requested a review from tnullJuly 15, 2026 14:01

@tnulltnull left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Comment threadsrc/lib.rs Outdated
_ = tokio::time::sleep(backoff) => {},
}

let undiscovered_lsps = rediscovery_ls.get_undiscovered_lsps();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex:

  • [P2] Background discovery can make add_liquidity_source fail spuriously. The /home/tnull/worktrees/ldk-node/pr-956-latest/src/lib.rs:760 can select a runtime-added LSP after it is inserted with supported_protocols: None but before /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/
    mod.rs:156. If the background task claims the pending-discovery slot first, the public call receives LiquidityRequestFailed and removes the otherwise valid LSP. Same-node discovery should be coalesced or initializing entries excluded from background sweeps.

Comment threadsrc/liquidity/mod.rs Outdated
node_id: PublicKey,
}

impl Drop for PendingDiscoveryGuard<'_> {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex

  • [P2] An old discovery guard can delete a newer request. The response handler removes the pending sender before waking its receiver, while /home/tnull/worktrees/ldk-node/pr-956-latest/src/liquidity/mod.rs:364 later removes whichever entry currently exists for that node. A second discovery
    can insert during that gap, only to have the first guard delete its sender. The guard should be disarmed after receiving its response or verify request identity before removal.

Comment threadsrc/lib.rs Outdated
backoff *= 2;
}

// periodically re-discover all configured LSPs to pick up newly

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused: we now have two APIs, right? Users can either add LSPs at build time, or at runtime. Both are not persisted, but the build-time one is likely 'persisted' in the user code. Why do we need to 'rediscover' these configured LSPs during the runtime - do we really expect our nodes to never restart and we hence require live protocol upgrades? I'm not sure we need to / should.

Add a background task that retries discovery for LSPs still undiscovered
after the startup batch, with exponential backoff (5s up to a 1h cap),
until every configured LSP is discovered. This recovers LSPs that were
unreachable at startup instead of leaving them permanently unusable.
Also coalesce concurrent discovery for the same LSP, so the retry task
and a runtime add_liquidity_source don't race and needlessly fail.
Look up trust_peer_0conf by node id via a protocol-independent helper
that does not depend on discovery
@Camillarhi
Camillarhiforce-pushed the multi-lsp-support-follow-up branch from 78c313f to 48104f7CompareJuly 21, 2026 01:06
@Camillarhi

Copy link
Copy Markdown
ContributorAuthor

Some comments that Codex had, mostly.

The cancellation safety part is now also addressed in #997.

Thanks, I reverted the cancellation-safety guard since #997 covers it. If #997 lands first, I'll rebase on top of it.

@Camillarhi
Camillarhi requested a review from tnullJuly 21, 2026 01:14
@tnull
tnull merged commit 08efb3a into lightningdevkit:mainJul 21, 2026
24 of 36 checks passed
This was referenced Jul 21, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@Camillarhi@ldk-reviews-bot@tnull