This library provides a set of C++ classes and functions for preprocessing sensor data. We split it out as a separate code base to enable reuse across multiple projects. Include this as a git submodule in your project to take advantage of its functionality.
The library exposes a Dart API through FFI (Foreign Function Interface). This allows Dart applications to call the C++ functions for sensor data preprocessing seamlessly.
In addition to the Dart API, the library also provides a Python package called senpy that provides acces to the C++ routines via Pybind11.
GitHub release wheels are built by .github/workflows/release-wheel.yml and attached to a release when it is published. The same workflow can be run manually to backfill an existing tag/release, such as v1.0.0 or v2.0.0.
Use the release asset URL directly when installing a precompiled wheel:
python -m pip install https://github.com/<owner>/<repo>/releases/download/v2.0.0/senpy-2.0.0-cp311-cp311-linux_x86_64.whlpip install git+https://github.com/<owner>/<repo>.git@v2.0.0 installs from source and will still compile the native extension locally.
compute_nustft needs the whole recording in memory. StreamingNUSTFT computes the same
coefficients from a live stream: push samples as they arrive, get each window back as soon as
the data passes its end, and never retain the samples themselves.
fromsenpyimportStreamingNUSTFTtransform=StreamingNUSTFT(
window_s=30.0, overlap_s=0.0, subwindow_s=1.0, # subwindow = one sensor packetsample_rate_hz=100.0, fmax=5.0, # report DC..5 Hz only
)
forpacket_t, packet_xinpackets:
forwindowintransform.push(packet_t, packet_x):
consume(window.center, window.magnitude())
tail=transform.flush() # the partly-filled final windowIt is exact, not an approximation. Against compute_nustft on the same samples the coefficients
agree to ~1e-13 relative (tests/test_streaming_nustft.py), for any chunking of the input and
with or without window overlap.
The transform is linear in the data and the subwindows partition the window, so
where
The contrast worth being clear about is with Bartlett/Welch. Averaging
Two parts of compute_nustft belong to the window rather than to any subwindow, and so are
deferred and reconstructed when the window closes:
The Hann taper. A subwindow cannot know its own taper weights until it knows where it sits.
On this frequency grid — spacing
so applying it after the fact is a three-tap convolution,
Mean removal. The window mean is unknown until the window closes, so each subwindow also
carries the transform of the constant 1. Detrending is then
with no second pass over the samples.
Each sample is touched once, at O(bins), however many windows it belongs to — so overlap is
nearly free, unlike the batch transform which re-spreads every sample per window. Memory is one
accumulator per open window plus the subwindows in flight; it does not grow with window length or
recording length. A narrow fmax is what makes the per-sample constant small: 100 Hz into a 5 Hz
band at 30 s windows costs about 150 000 multiply-accumulates per second of stream.
- Ordering. Timestamps must be non-decreasing over the object's life. A sample belonging to a
subwindow the stream has already passed cannot be folded in;
dropped_samplescounts those. - Window grid.
origin_sanchors it, and window 0 is the earliest — nothing before the origin is reported. Pass the first timestamp to reproducecompute_nustft's alignment, or a fixed epoch to keep window indices meaningful across sessions and processes. - Divisibility. The window and the hop must be whole multiples of
subwindow_s, so that no subwindow straddles a window edge; one that did could not be shared by the windows either side. - Sample rate. Supplied rather than measured: it sets the magnitude scale and the grid size.
compute_nustfttakes the median spacing over the whole recording, which a stream cannot see. - The trailing window.
pushreports only windows the stream has passed the end of — all a live stream can honestly say.compute_nustftknows where the recording stops and also emits a final window ending within one sample period of the last timestamp; that one comes out offlush().compute_nustft_streamingapplies this rule for you and is the function to compare the two paths with. - The Nyquist bin (present only when
fmaxis unset) is the true$+N/2$ coefficient.compute_nustftreports its conjugate there, an artefact of reading that bin out of the aliased FINUFFT mode. Magnitudes are identical. - Timestamp precision. Absolute unix seconds in float64 resolve to about half a microsecond, which is a ~1e-5 relative phase error at the top of a 5 Hz band. Pass times relative to a recent origin when sub-microsecond timing matters.
The recombination identity and its resolution argument are developed in
autofish-jax/mobile/litert_spike/stft_recombination.tex. The first production user is
FoundryWhoopAndroid, which computes 30 s sleep-staging features one strap packet at a time
because it deletes the raw samples after upload; its Kotlin implementation and this one agree to
float32 storage precision on the same fixture.
The regular senpy API remains NumPy/C++ based. For a JAX-native NUFFT that
keeps sample arrays on the active JAX device, install jax-finufft with the
JAX build appropriate for the machine:
# First install JAX with its CUDA support using JAX's installation guidance.# Then, from this repository's senpy/ directory:
python -m pip install '.[jax]'importjaximportjax.numpyasjnpfromsenpyimportjax_backendassenpy_jaxtimestamps=jnp.arange(3_000, dtype=jnp.float32) /50.0signal=jnp.sin(2*jnp.pi*3.0*timestamps)
result=senpy_jax.compute_nustft(
timestamps, signal, window_s=8.0, overlap_s=4.0, target_fs=16.0
)
print(jax.devices(), result.coefficients.shape)senpy.jax_backend uses jax-finufft's type-1 transform; a CUDA-enabled
jax-finufft build dispatches it to cuFINUFFT. It returns JAX arrays rather
than senpy.api.NUSTFTResult, so subsequent JAX work stays device-resident.
The GPU default is eps=1e-6; enable JAX x64 before importing JAX if the
application requires float64 precision. Absolute epoch timestamps are safe to
pass as NumPy arrays -- they are centered on the first sample in float64 before
reaching the device. If you build the timestamp array with JAX yourself, either
enable x64 first or make the values relative to the first sample; float32 cannot
resolve millisecond spacing at epoch magnitude, and compute_nustft rejects
such an array rather than returning a wrongly scaled result.
For high-throughput three-axis work across recordings, pre-pack ragged windows into a small set of static shapes, then run each batch on the JAX device:
fromsenpyimportjax_backendassenpy_jax# Each sample array is shaped [N, 3] for x/y/z. The packer only discovers and# pads windows; it does not import JAX or execute a transform.batches=senpy_jax.pack_nustft_window_batches(
recordings, window_s=8.0, overlap_s=4.0, batch_size=128, ts_unit="s"
)
forbatchinbatches:
coefficients=senpy_jax.compute_nustft_window_batch(
batch.points,
batch.signals,
batch.valid,
nfft_padded=batch.nfft_padded,
median_fs=batch.median_fs,
)
real_coefficients=coefficients[batch.row_valid] # [windows, 3, freqs]recording_indices, window_indices, and times in each batch map valid
output rows back to the input order. Batch sizes remain a hardware-specific
throughput setting: measure with block_until_ready() and a CUDA profiler
before claiming GPU saturation.