Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Sensor Preprocessing C++ Library

This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.

Dart API

The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.

Python API

In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.

Precompiled Python wheels

GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.

Use the release asset URL directly when installing a precompiled wheel:

python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whl

pip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.

Streaming NUSTFT

compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same coefficients from a live stream: push samples as they arrive, get each window back as soon as the data passes its end, and never retain the samples themselves.

fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final window

It is exact, not an approximation. Against compute_nustft on the same samples the coefficients agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and with or without window overlap.

Why it is exact

The transform is linear in the data and the subwindows partition the window, so

$$X_w(f) ;=; \sum_m e^{2\pi i f d_m}, S_m(f)$$

where $S_m$ is the transform of subwindow $m$ about its own origin and $d_m$ is that origin's offset into the window. This is decimation-in-time for nonuniformly sampled data: a long transform is a phase-weighted sum of short ones, with nothing lost.

The contrast worth being clear about is with Bartlett/Welch. Averaging $|S_m|^2$ over subwindows pins the frequency resolution at the subwindow's$1/T_w$ — averaging buys variance reduction, never resolution. Keeping the complex $S_m$and their offsets recovers the full window's transform, hence the full window's resolution. Coherence is the whole trick.

Two parts of compute_nustft belong to the window rather than to any subwindow, and so are deferred and reconstructed when the window closes:

The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits. On this frequency grid — spacing $1/T$, because that is the grid the window's own transform lives on — the taper is exactly one bin wide:

$$h(\tau) = \tfrac12 - \tfrac14 e^{2\pi i \tau} - \tfrac14 e^{-2\pi i \tau}$$

so applying it after the fact is a three-tap convolution, $0.5,D(k) - 0.25,D(k-1) - 0.25,D(k+1)$. That is why one bin above the reported band is carried internally; bin $-1$ needs no storage, being the conjugate of bin $+1$ for real input.

Mean removal. The window mean is unknown until the window closes, so each subwindow also carries the transform of the constant 1. Detrending is then $D(k) = X(k) - \bar{x},\mathrm{ones}(k)$. The same accumulator pays for the scale factor too: expanding $h^2$ gives $0.375 - 0.5\cos 2\pi\tau + 0.125\cos 4\pi\tau$, so

$$\sum_j h(\tau_j)^2 ;=; 0.375,N - 0.5,\mathrm{Re},\mathrm{ones}(1) + 0.125,\mathrm{Re},\mathrm{ones}(2)$$

with no second pass over the samples.

Cost

Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one accumulator per open window plus the subwindows in flight; it does not grow with window length or recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.

Contract and differences from compute_nustft

  • Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a subwindow the stream has already passed cannot be folded in; dropped_samples counts those.
  • Window grid.origin_s anchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproduce compute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes.
  • Divisibility. The window and the hop must be whole multiples of subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side.
  • Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size. compute_nustft takes the median spacing over the whole recording, which a stream cannot see.
  • The trailing window.push reports only windows the stream has passed the end of — all a live stream can honestly say. compute_nustft knows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out of flush(). compute_nustft_streaming applies this rule for you and is the function to compare the two paths with.
  • The Nyquist bin (present only when fmax is unset) is the true $+N/2$ coefficient. compute_nustft reports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical.
  • Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.

Where this came from

The recombination identity and its resolution argument are developed in autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time because it deletes the raw samples after upload; its Kotlin implementation and this one agree to float32 storage precision on the same fixture.

JAX / NVIDIA CUDA NUFFT

The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that keeps sample arrays on the active JAX device, install jax-finufft with the JAX build appropriate for the machine:

# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'
importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)

senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident. The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the application requires float64 precision. Absolute epoch timestamps are safe to pass as NumPy arrays -- they are centered on the first sample in float64 before reaching the device. If you build the timestamp array with JAX yourself, either enable x64 first or make the values relative to the first sample; float32 cannot resolve millisecond spacing at epoch magnitude, and compute_nustft rejects such an array rather than returning a wrongly scaled result.

For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:

fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]

recording_indices, window_indices, and times in each batch map valid output rows back to the input order. Batch sizes remain a hardware-specific throughput setting: measure with block_until_ready() and a CUDA profiler before claiming GPU saturation.

About

C++ library with Dart and Python bindings, implementing accelerometer preprocessing for machine learning models.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages