Skip to content

Single-precision (fp32) build support - #5033

Draft
hardik-corintis wants to merge 99 commits into
firedrakeproject:mainfrom
Corintis:fp32-support
Draft

Single-precision (fp32) build support#5033
hardik-corintis wants to merge 99 commits into
firedrakeproject:mainfrom
Corintis:fp32-support

Conversation

@hardik-corintis

@hardik-corintis hardik-corintis commented Apr 15, 2026

Copy link
Copy Markdown
Contributor

fixes #3040

Description

Adds single-precision (fp32) build support. Firedrake can now run on a PETSc installation compiled with --with-precision=single. The approach mirrors complex mode: precision is detected at import time from PETSc's build variables and flows through from there.

AI disclosure: Parts of this PR were developed with assistance from Claude (Anthropic). All changes have been reviewed, tested locally, and are fully understood by the author.

Prerequisite

Requires https://gitlab.com/petsc/petsc/-/merge_requests/9272 (petsc4py: handle PETSC_DOUBLE in DMSwarm.getField). Without it, DMSwarm.getField() raises AssertionError on fp32 builds because PETSC_DOUBLE is not mapped to a numpy dtype in the single-precision case where PETSC_REAL != PETSC_DOUBLE.

Changes

scripts/firedrake-configure

  • Adds --arch single / ScalarType.SINGLE; passes --with-precision=single to PETSc configure; excludes fftw and suitesparse (no fp32 support in those libraries)

tsfc/parameters.py, tsfc/loopy.py, tsfc/kernel_interface/common.py, tsfc/ufl_utils.py

  • scalar_type / scalar_type_c derived from PETSc precision at import time; constant initializers cast to the kernel scalar dtype

Core (evaluate.h, locate.c, pointquery_utils.py, pointeval_utils.py, mg/kernels.py)

  • Replace hardcoded double / int with PetscReal / PetscInt in generated C code
  • Convergence epsilon in point query tightened to 1e-6 in fp32 mode vs 1e-12 in fp64, to stay within single-precision range

firedrake/mesh.py, firedrake/utility_meshes.py

  • Vertex coordinates and reference-cell distances use PETSc.RealType; physical coordinate arrays for rtree and DMSwarmPIC_coor remain float64 (required by the rtree C API and PETSc swarm internals)

firedrake/function.py

  • Point evaluation coerces coordinates to float64 regardless of ScalarType, for geometric robustness in cell location

firedrake/assemble.py, firedrake/functionspaceimpl.py, pyop2/codegen/builder.py

  • Replace dtype=int with dtype=IntType in numpy.prod and array allocation calls

firedrake/utils.py

  • Adds single_mode boolean flag (mirrors complex_mode)

Tests

  • Adds @pytest.mark.skipsingle marker for tests incompatible with fp32
  • tests/firedrake/conftest.py: registers the marker and wires it up
  • Skips a small set of tests that require double-precision accuracy (test_locate_cell, test_interpolate_cross_mesh[extrudedcube], test_parallel_high_order_location)

.github/workflows/core.yml

  • Adds single to the CI matrix alongside default and complex

Known limitations

test_parallel_high_order_location is skipped in fp32: high-order cell location in a warped mesh requires double-precision accuracy that fp32 cannot provide at tolerance=0.0001.

@hardik-corintis
hardik-corintis marked this pull request as draft April 15, 2026 14:13
@connorjward

Copy link
Copy Markdown
Contributor

Thanks for this.

adds --arch single

What about single+complex? Isn't that a valid configuration?

PETSc version bump (v3.24.5 → v3.25.0)

This isn't necessary. That's all going to be taken care of when I release the next major version in the next 24 hours.

Fix needed upstream in petsc4py: if ctype == PETSC_DOUBLE: typenum = NPY_DOUBLE in petsc4py/PETSc/DMSwarm.pyx.

Can you get this fixed upstream? Clearly Claude already knows what to do.

@hardik-corintis
hardik-corintis force-pushed the fp32-support branch 2 times, most recently from 0dc7472 to dd517a9 Compare May 9, 2026 23:10
@hardik-corintis
hardik-corintis marked this pull request as ready for review May 10, 2026 00:18
@hardik-corintis

Copy link
Copy Markdown
Contributor Author

Thanks for this.

adds --arch single

What about single+complex? Isn't that a valid configuration?

Added ScalarType.SINGLE_COMPLEX (--arch single-complex) to firedrake-configure for all four platform targets: single_mode (from PETSC_PRECISION) and complex_mode (from PETSC_SCALAR) are detected independently, and ScalarType already resolves to numpy.complex64 for a single+complex build.

PETSc version bump (v3.24.5 → v3.25.0)

This isn't necessary. That's all going to be taken care of when I release the next major version in the next 24 hours.

Removed from the PR description — thanks.

Fix needed upstream in petsc4py: if ctype == PETSC_DOUBLE: typenum = NPY_DOUBLE in petsc4py/PETSc/DMSwarm.pyx.

Can you get this fixed upstream? Clearly Claude already knows what to do.

Opened a PR here: https://gitlab.com/petsc/petsc/-/merge_requests/9272

@connorjward connorjward left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've just looked at the CI/install related side of this (at a glance everything else seems pretty good).

We just have to be careful about adding new test builds to Firedrake. I will try and figure out a solution soon.

Comment thread .github/workflows/core.yml Outdated
Comment thread scripts/firedrake-configure Outdated
@connorjward

Copy link
Copy Markdown
Contributor

We just have to be careful about adding new test builds to Firedrake. I will try and figure out a solution soon.

I have a solution for this in #5117 (merged). Once we update main (#5134) then you can use it in this pull request. You basically need to follow the same procedure that we have for testing complex or CUDA builds.

- mesh.py: clarify why PETSc.RealType is correct for plex_from_cell_list
- mg/kernels.py: fix to_reference_coords_kernel signature (PetscScalar *X -> %(RealType)s *X)
  and add RealType to the template substitution dict
- evaluate.h: comment explaining why int/double is intentional for evaluate()
- test_stokes_mini.py: replace fp32 solver workaround with @pytest.mark.skipsingle;
  fieldsplit path needs a dedicated fp32 test
- conftest.py: remove trailing blank line
# Conflicts:
#	firedrake/function.py
#	firedrake/mesh.py
#	scripts/firedrake-configure
# Conflicts:
#	.github/workflows/core.yml
#	.github/workflows/pr.yml
#	.github/workflows/push.yml
#	firedrake/evaluate.h
#	firedrake/function.py
#	firedrake/locate.c
#	firedrake/mesh.py
#	firedrake/nullspace.py
#	firedrake/pointeval_utils.py
#	scripts/firedrake-configure
#	tests/firedrake/adjoint/test_transformed_functional.py
#	tests/firedrake/multigrid/test_adaptive_multigrid.py
#	tests/firedrake/regression/test_bddc.py
#	tests/firedrake/vertexonly/test_vertex_only_mesh_generation.py
Threads dtype=RealType (or the numpy-derived real dtype from scalar_type
in TSFC) through every as_fiat_cell/create_element call site in
firedrake and tsfc, using FIAT PR firedrakeproject#260 and FInAT's new dtype support
(firedrakeproject/fiat#263). Removes the finat.element_factory.ufc_cell
monkeypatch in firedrake/utils.py, which is no longer needed now that
callers declare their working precision explicitly instead of relying
on a globally-overridden default.

Also fixes petsc_sparse's hardcoded rtol=1E-10 drop tolerance
(firedrake/preconditioners/fdm.py), which is too tight for single
precision: reference-cell geometry feeding into these matrices is only
accurate to float32 round-off (~1e-7), so entries that should be
exactly zero (e.g. in a Nedelec discrete gradient matrix) were leaking
into the sparsity pattern above the fp64-tuned threshold, corrupting
the exact combinatorial structure PETSc's PCBDDCNedelecSupport expects
and causing test_bddc_aij_simplex[N1curl-3-False] to fail.
test_fieldsplit: h *= 8.0 (was 10.0) for the Taylor test perturbation.
test_transformed_functional_poisson_tao_nls: widen the fp32 error_norm
bound to 2e-1 (was 7e-2) and correct the comment -- round-off shifts
which critical point tao_nls converges to, not a solver tolerance
issue.
@pbrubeck

pbrubeck commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Some of the tests could get a global tolerance per module. There's no need to add repeated logic each time a tolerance is used

@hardik-corintis
hardik-corintis marked this pull request as draft August 7, 2026 15:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci:single Run the test suite in single precision

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support precisions other than double

4 participants