portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum - #6

Closed
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision
Closed

portable: accumulate in fp32 for Half/BFloat16 in grid_sampler_2d and sum#6
jgibson2 wants to merge 1 commit into
mainfrom
jgibson/portable-fp16-precision

Conversation

@jgibson2

Copy link
Copy Markdown
Collaborator

Summary

Both grid_sampler_2d.out (bilinear) and sum.IntList_out in the portable kernels previously did all interior arithmetic in the input tensor's dtype. For fp16 inputs that's a material precision problem, not just FP rounding noise:

grid_sampler_2d.out (bilinear)

Interpolation weights are derived from subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands are close integer values. In fp16, subtracting two close values of the form `N.xxx` and `N+1.xxx` with only ~10 bits of mantissa destroys most of the precision — classic catastrophic cancellation. The downstream weighted sum then accumulates further error in fp16.

sum.IntList_out

The fast path (innermost contiguous dim):

```cpp
CTYPE acc = 0;
for (int64_t j = 0; j < reduce_size; j++) {
acc += row[j]; // fp16 += fp16 over many elements
}
```

Over a reduction of more than ~100 fp16 values of similar magnitude, cumulative error grows to the same order as the reduction result itself. The slow path (via `MapReduceOverDimListPlan`) had the same issue because its `CTYPE_OUT` template parameter was used for the accumulator.

Fix

In each file: an `AccType` / `SumAccType` trait that maps `Half` and `BFloat16` to `float` and leaves every other dtype unchanged. The trait is used for the internal accumulator, intermediate coordinate / weight computation, and the single cast back to the output dtype at store time. Loads and stores remain in the tensor's dtype — only the inner arithmetic is promoted.

```cpp
template
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```

Measured effects

On a unit test comparing "fp16 input → fp16 output" against "run in fp32 and cast at the end" for the shapes actually exercised by the polycam depth model:

OpBeforeAfter
`grid_sampler_2d` bilinear, fp16, random interior gridmax_abs ≈ 0.10max_abs = 0
`sum.IntList_out`, fp16, reduce over 256-element innermostmax_abs ≈ 0.14max_abs = 7.6e-6

Non-effects

  • fp32 / Int / any other dtype: byte-identical output. `AccType` is `T` for non-half types, so the generated code is unchanged.
  • Perf: overhead is a handful of fp16↔fp32 conversions per output element (cheap on any modern CPU with NEON / AVX). Not measurable at the op level.
  • No public API change. Signatures unchanged, no new flags, no new allocations.

Test plan

  • Builds cleanly on Android arm64 and host (Apple Clang 21).
  • Verified numerically via a standalone harness that runs each kernel with matched fp32 / fp16 inputs against an fp32 reference (run-in-fp32 then downcast). All shapes tested pass within fp16 ULP; fp32 paths bit-identical.
  • End-to-end model validation (to be done by the PR author on a trained polycam depth model — not strictly a blocker for landing since behavior on fp32 is unchanged).

Candidate for upstream at some point — this is a general correctness improvement, not Polycam-specific.

🤖 Generated with Claude Code

… sum
Both kernels previously performed all interior arithmetic in the input
tensor's dtype. For half-precision inputs that's a material precision bug:
* grid_sampler_2d.out (bilinear): interpolation weights are derived from
subtractions like `(ix_se - ix)` and `(iy_se - iy)` where both operands
are close integer values. In fp16 that's catastrophic cancellation —
the result has only a handful of significant bits. The weighted sum
then further accumulates error in fp16.
* sum.IntList_out: the fast path (innermost contiguous dim) uses a scalar
accumulator of input dtype: `CTYPE acc = 0; acc += row[j]`. Over a
reduction of >100 fp16 values, cumulative error reaches the same order
of magnitude as the reduction itself. Observed empirically on a real
model: sums of 256 fp16 values drift by up to ~0.14 absolute relative
to the fp32 reference. The slow path (MapReduceOverDimListPlan) had
the same issue via its CTYPE_OUT template parameter.
Fix in both files: introduce a simple `AccType<CTYPE>` / `SumAccType<CTYPE>`
trait that maps Half and BFloat16 to `float` and leaves all other types
unchanged. Use it for the internal accumulator, intermediate coordinate /
weight computation, and the single cast back to output dtype at store
time. Loads and stores remain in the tensor's dtype — only the inner
arithmetic is promoted.
Effects:
* fp32 / Int / other dtypes: byte-identical output (AccType is a no-op).
* fp16 / BFloat16: substantially tighter agreement with fp32 reference.
On a unit test exercising the actually-hot shapes in our depth model,
`max_abs` between "fp16 input → fp16 output" and "run in fp32 and cast
at the end" drops from ~0.14 to 7.6e-6 for sum and from ~0.1 to 0
for grid_sampler_2d bilinear.
* Perf: the overhead is a handful of fp16↔fp32 conversions per output
element. Not measurable at the op level — still well within the
scalar-portable-kernel cost envelope for Half inputs.
No public API change. No behavioral change for fp32 workloads.
@jgibson2

Copy link
Copy Markdown
CollaboratorAuthor

Superseded — the grid_sampler portion is now upstream at pytorch#19117, and the sum portion is dropped (not needed for the polycam use case once the NEON sum kernel is in place).

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jgibson2