Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Free holding cell on monitor-updating-restored when there's no upd by TheBlueMatt · Pull Request #755 · lightningdevkit/rust-lightning · GitHub
Skip to content

Free holding cell on monitor-updating-restored when there's no upd - #755

Closed
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd
Closed

Free holding cell on monitor-updating-restored when there's no upd#755
TheBlueMatt wants to merge 1 commit into
lightningdevkit:mainfrom
TheBlueMatt:2020-11-holding-cell-clear-mon-upd

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.

Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).

Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.

I didn't bother writing a test cause I'm not entirely sure this is the right approach - mostly I'd like concept feedback.

If there is no pending channel update messages when monitor updating
is restored (though there may be an RAA to send), and we're
connected to our peer and not awaiting a remote RAA, we need to
free anything in our holding cell.
Without this, chanmon_fail_consistency was able to find a stuck
condition where we sit on an HTLC failure in our holding cell and
don't ever handle it (at least until we have other actions to take
which empty the holding cell).
Still, this approach sucks - it introduces reentrancy in a
particularly dangerous form:
a) we re-enter user code around monitor updates while being called
from user code around monitor updates, making deadlocks very
likely (in fact, our current tests have a bug here!),
b) the re-entrancy only occurs in a very rare case, making it
likely users will not hit it in testing, only deadlocking in
production.
I'm not entirely sure what the alternative is, however - we could
move to a world where we poll for holding cell events that can be
freed on our 1-minute-timer, but we still have a super rare
reentrancy case, just in timer_chan_freshness_every_min() instead.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Because I'm not sure if this is the right fix, and the case should be incredibly rare, and its not really a critical bug, per se, and this is a more involved change than everything else that's left, I'm not going to bother tagging this 0.0.12 and suggest we hold off on fixing it.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Oh, I should mention this was found by the updates in #753.

@codecov

codecovBot commented Nov 19, 2020

Copy link
Copy Markdown

Codecov Report

Merging #755 (07fb28e) into main (4e82003) will decrease coverage by 0.06%.
The diff coverage is 93.18%.

Impacted file tree graph

@@ Coverage Diff @@## main #755 +/- ##
==========================================
- Coverage 91.45% 91.38% -0.07% 
==========================================
Files 37 37 Lines 22249 22263 +14 ==========================================
- Hits 20347 20346 -1 - Misses 1902 1917 +15 
Impacted FilesCoverage Δ
lightning/src/ln/channel.rs87.37% <81.25%> (-0.07%)⬇️
lightning/src/ln/channelmanager.rs85.48% <100.00%> (+0.11%)⬆️
lightning/src/ln/functional_tests.rs96.92% <0.00%> (-0.25%)⬇️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4e82003...f35c8bf. Read the comment docs.

@valentinewallace

Copy link
Copy Markdown
Contributor

Could I see the fuzz failure that this is solving?

I'm having a hard time wrapping my mind around #756 and I think it's because it seems to be solving multiple distinct problems that could be separated out. So, gonna try to understand them one by one.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

@valentinewallace

valentinewallace commented Feb 10, 2021

Copy link
Copy Markdown
Contributor

This input to chanmon_consistency fails on current upstream.

3c10143411341434100d5454543033333038345454545454545454545454
5454545454545454545454545454110015542d3454545411151154543154
541cff14

For me this fuzz input^ fails with the same error on main and this PR's commit. Is that expected?

If this PR can stand on its own -- and it's not too inconvenient -- I'd rather reopen, review + get this mergeable separately from addressing the other commits in #756.

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Right, honestly its been a while and I don't remember exactly the contours of what is/isn't in this PR, but that fuzz input does pass on #756. Most of the other commits on 756 are just merging to codepaths that were near-identical that should now be identical and then fixing a few miner further bugs in the fuzztarget. I'm happy to try to dig deeper into it, but I'm not sure just this version is really an option, so much as an intended demonstration of an idea seeking concept acks.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@TheBlueMatt@valentinewallace