Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); UniformFloat: allow inclusion of high in all cases by dhardy · Pull Request #1462 · rust-random/rand · GitHub
Skip to content

UniformFloat: allow inclusion of high in all cases - #1462

Merged
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats
Jul 16, 2024
Merged

UniformFloat: allow inclusion of high in all cases#1462
dhardy merged 8 commits into
rust-random:masterfrom
dhardy:closed-range-floats

Conversation

@dhardy

@dhardydhardy commented Jun 24, 2024

Copy link
Copy Markdown
Member
  • Added a CHANGELOG.md entry

Summary

Fix#1299 by removing logic specific to ensuring that we emulate a closed range by excluding high from the result.

Motivation

Fix#1299.

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

In general, one can always expect that floating-point numbers may round up/down to the nearest representable value, hence this change is not expected to affect many users.

sample_single was changed (see below) as a simplification, and to match the new behaviour of new.

Details

Specific to floats (i.e. UniformFloat),

  • new sets scale = high - low (unchanged) then ensures that low + scale * max_rand <= high (previously: ensure < high)
  • new_inclusive sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)
  • sample_single is now equivalent to sample_single_inclusive (previously: use rejection sampling to ensure sample < high)
  • sample_single_inclusive is unchanged: it yields low + (high-low) * Standard.samgle(rng)

Note: new attempts to ensure that high may be yielded; sample_single_inclusive does not. This has not changed, but is a little surprising.

@dhardy

Copy link
Copy Markdown
MemberAuthor

I should add that there is still a loop in new and new_inclusive, but I don't think it can be a problem any more. This would require that low + (high - low) / (1 - ε) * (1 - ε) rounds to a value greater than high, but additionally that there are many representable values between scale = (high - low) / (1 - ε) and the next smallest value which will round down to high. Specifically, we know that high is representable.

@sicking@WarrenWeckesser

@vks

vks commented Jun 27, 2024

Copy link
Copy Markdown
Contributor

Additional motivation: floating-point types are approximations of continuous (real) numbers. Although one might expect that rng.gen_range(0f32..1f32) is strictly less than 1f32, it can also be seen as an approximation of the real range [0, 1) which includes values whose nearest approximation under f32 is 1f32.

I don't like this argument, because when calculating probabilities with real numbers, there is no difference between [0, 1) and [0, 1] when integrating the PDF. Therefore, the distinction only makes sense for floats. Also, when using something advertised as [0, 1), I would want to be able to rely on it, because otherwise I have to take care of the corner cases with rejection sampling when e.g. taking logs. If we cannot guarantee it, we might as well deprecate [0, 1) and only offer (0, 1).

I agree we should fix #1462.

new sets scale = (high - low) / (1 - ε) then ensures that low + scale * max_rand <= high (unchanged)

I think you meant new_inclusive?

@dhardy

dhardy commented Jun 27, 2024

Copy link
Copy Markdown
MemberAuthor

I don't like this argument ...

With real numbers, both [0, 1] and [0, 1) have a range of 1, but the latter does not contain 1. With floats, it's pretty easy to lose precision — e.g. (1 - ε/2) + 1 → 2 — so does it make sense for us to take great care here? Will many users actually care that we do?

new and new_inclusive are part of trait UniformSampler and important to integers.

Meanwhile, we still provide Standard, OpenClosed01 and Open01 (but not Closed01). None of those are affected here.

I agree we should fix #1462.

I think you meant #1299?

@dhardy
dhardy requested a review from TheIronBornJuly 8, 2024 08:46
@dhardydhardy added D-review Do: needs review X-discussion labels Jul 8, 2024

@TheIronBornTheIronBorn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks

@dhardy

Copy link
Copy Markdown
MemberAuthor

I am now reasonably convinced that there are no (low, high) pairs of floating-point values such that more than one iteration of the loop in UniformFloat::new_bounded is required. Proof follows (and focusses on new_inclusive, which is the worst case).

Definition: ε is the smallest value such that 1 + ε is a representable value greater than 1.

Assumption: low, high are finite floating-point values.
Assumption: low ≤ high.
Assumption: (high - low) / (1 - ε) is finite.
(These assumptions are tested in new_inclusive.)

new_inclusive calculates scale = (high - low) / (1 - ε), then decreases to the next smallest value while scale × (1 - ε) > high.

Case 1: high - low rounds up. This is possible when low and high have opposite signs such that the result is larger than the absolute value of either low or high. An example is low = -1, high = 3ε/2 resulting in high - low → 1 + 2ε. Here, the exponent of the result is one larger than the maximum exponent of low and high. Since this is a rounding error, it is at worst one value too large.

Case 2: x / (1 - ε) rounds up such that x / (1 - ε) * (1 - ε) > x (where x = high - low). Note that floating-point numbers have format (-1)^sign * (1 + significand) * 2^exponent where 0 ≤ significant < 1; division by 1 - ε effectively increments the significand by 1 or 2 (or 0 in the case of sub-normals). In some cases this rounds up; I believe the worst (relative) case of this is (2 - ε) / (1 - ε) → 2 + 2ε such that the result, when multiplied by 1 - ε, is 2 (one representable value larger than 2 - ε).

In fact, I believe case 2 can only happen where the division causes a loss of precision (i.e. increases the exponent since x is very nearly the largest representable value with its exponent), and thus that the subtraction step (case 1) cannot also round up.

(We could thus potentially replace the loop with a single-step if, and I think we could remove it entirely for new (exclusives). These are only minor optimisations and thus excluded here.)

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

D-reviewDo: needs review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uniform Generator hangs for certain limits.

3 participants

@dhardy@vks@TheIronBorn