Skip to content

Add the vendor VFIO vGPU device backend - #364

Open
yummybomb wants to merge 24 commits into
hypeship/hypervisor-livenessfrom
hypeship/vendor-vfio-backend
Open

Add the vendor VFIO vGPU device backend#364
yummybomb wants to merge 24 commits into
hypeship/hypervisor-livenessfrom
hypeship/vendor-vfio-backend

Conversation

@yummybomb

@yummybombyummybomb commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Summary

Layer 2 of the vendor VFIO vGPU stack (generalize-vgpu-devicehypervisor-livenessthisvendor-vfio-vgpu). Self-contained in lib/devices + lib/resources; nothing in the instance lifecycle calls it yet (that's the top layer).

Linux 6.8 hosts with NVIDIA R580 drop the mdev interface: vGPUs are assigned by writing a type ID to a VF's nvidia/current_vgpu_type and passed to QEMU as a plain VFIO PCI device. This adds that backend behind the framework dispatch introduced in #322:

  • Discovery & placement — profile discovery from the capacity-dependent creatable_vgpu_types catalogs, least-loaded-GPU VF selection, create/verify/rollback. Profile availability counts free VFs currently advertising each type as a best-effort snapshot; creating one assignment may change sibling catalogs.
  • Release guards — unlike mdev (fresh UUID per assignment), vendor VFIO reuses the same VF path across assignments, so a stale release could clear a later owner's vGPU. Release is guarded by an in-process owner map (covers the window before QEMU opens the device) and an open-VFIO-handle scan (refuses to clear a VF a running VM holds).
  • Reconciliation — clears orphaned assignments on startup, skips VFs in the caller-supplied protected set, and fails closed (skips vendor VFIO entirely) when that set is unavailable, while mdev reconciliation still runs.
  • Per-VF degradation — one unreadable VF is skipped with a warning instead of failing discovery or profile listing wholesale, so a flaky sysfs read cannot blank the host's advertised GPU capacity. Only when no VF is readable does discovery fail, so a wholesale outage cannot demote a vGPU host to passthrough while assignments exist. The open-handle probe stays strict on purpose: it authorizes clearing a reused VF path, so an incomplete scan fails the release rather than risk a false "not in use".
  • Integration test — branched by discovered framework, extended to cover release on stop and reacquisition on start.

Testing

  • go build ./..., go vet clean
  • go test -race ./lib/devices/ ./lib/resources/ pass

Note

High Risk
Touches GPU assignment, sysfs writes, and resource admission: a discovery or release bug can leak VFs, steal another instance’s assignment, or hide/expose GPU capacity incorrectly. Create is still gated, which limits production blast radius until the next layer.

Overview
Adds a Linux vendor VFIO vGPU backend for hosts that assign profiles via nvidia/current_vgpu_type instead of mdev (NVIDIA R580 / kernel 6.8+). Discovery prefers mdev, then vendor VFIO, and only then passthrough—so a failed vGPU probe no longer demotes a vGPU host to passthrough.

The backend discovers VFs, lists creatable_vgpu_types as a best-effort availability snapshot, places on the least-loaded GPU, and create/verify/rollback by writing the type ID. Release is guarded by an in-process owner map plus a strict open-VFIO-handle scan (because VF paths are reused). Reconciliation clears orphans unless they are protected or still held open. Unreadable VFs are skipped rather than wiping advertised capacity.

CreateVGPU still rejects vendor VFIO until instance-lifecycle integration lands. Resource status, OpenAPI profile available wording, GPU.md, and the vGPU integration test (stop/start reacquire; skip vendor VFIO for now) are updated to match.

Reviewed by Cursor Bugbot for commit 6f98055. Bugbot is set up for automated code reviews on this repo. Configure here.

Comment threadlib/devices/vendor_vfio_linux.go Outdated
Comment threadlib/resources/gpu.go
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 2bb8e86 to 7fc3b49CompareAugust 6, 2026 19:26
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 7fc3b49 to f661e63CompareAugust 6, 2026 19:40
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from f661e63 to d1207d0CompareAugust 7, 2026 14:02
Comment threadlib/resources/gpu.go
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from d1207d0 to 5b47670CompareAugust 7, 2026 15:04
Comment threadlib/devices/vendor_vfio_linux.go
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 5b47670 to ffbf8a0CompareAugust 7, 2026 20:52
Comment threadlib/devices/vendor_vfio_linux.go
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from ffbf8a0 to 78801e7CompareAugust 8, 2026 01:05
Comment threadintegration/vgpu_test.go Outdated
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 78801e7 to b4651f3CompareAugust 9, 2026 06:26
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from b4651f3 to d2bf3faCompareAugust 9, 2026 07:52
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from d2bf3fa to 495ab62CompareAugust 9, 2026 23:27
@yummybomb

Copy link
Copy Markdown
ContributorAuthor

added 7f5233f: report GPUProfile.Available as the count of free VFs advertising the profile type (creatable-instance units), matching mdev's summed available_instances, the OpenAPI description, and the integration test's decrement assertion. note the gpuProfileSlots metric steps up on vendor VFIO hosts (per-GPU → per-free-VF units). go test ./lib/devices ./lib/resources green; the hardware integration test was not run locally.

@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 7f5233f to 0baacb3CompareAugust 10, 2026 07:26
@github-actions

github-actionsBot commented Aug 10, 2026

Copy link
Copy Markdown
-->

✱ stlc build

gocode · compare

Your SDK build was successful.

generate ✅bootstrap ✅format ✅

116 files generated at f6146e4 (pushed)

go get github.com/kernel/hypeman-go-staging@f6146e4cb17b257dd8cdbc01ad7fc6259cce0017
pythoncode · compare

Your SDK build was successful.

generate ✅bootstrap ✅format ✅

230 files generated at 3c815a0 (pushed)

typescriptcode · compare

Your SDK build was successful.

generate ✅bootstrap ✅format ✅

138 files generated at 2cbb81b (pushed)

Diagnostics: ❗ 0 new / 1 total error, 💡 0 new / 5 total note
LevelCodeMessageTargets
Build metadata
Buildbd_76BcNdZp-glacial-scale
Timestamp2026-08-21T20:51:21.988Z
stlc8413509
Spec hasha37dc5305160
Config hash55e15f6f4434

This comment is auto-generated by stlc and is kept up to date as you push.
If you push new commits, re-run this workflow to update this comment.
Last updated: 2026-08-21 20:51:46 UTC

@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 1f1f24f to d28e3d0CompareAugust 20, 2026 22:04

@cursorcursorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit d28e3d0. Configure here.

Comment threadlib/devices/vendor_vfio_linux.go
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from e44df0b to 0ddc917CompareAugust 21, 2026 20:44
Linux 6.8 hosts with NVIDIA R580 drop the mdev interface: vGPUs are
assigned by writing a type ID to a VF's nvidia/current_vgpu_type and
passed to QEMU as a plain VFIO PCI device. Add a vendor VFIO backend
behind the existing framework dispatch: profile discovery from the
capacity-dependent creatable catalogs, least-loaded VF placement,
create/verify/rollback, and release.
Because the same VF path is reused across assignments (unlike mdev
UUIDs), release is guarded: an in-process owner map covers the window
before QEMU opens the device, and an open-VFIO-handle scan refuses to
clear a VF a running VM still holds. Reconciliation clears orphaned
assignments on startup, skipping VFs protected by the caller and
failing closed when the protected set is unavailable.
Branch the vGPU integration test by discovered framework and extend it
to cover release on stop and reacquisition on start.
Sort GPUs with unaccountable load last instead of rejecting placement, and stop reporting passthrough capacity when vGPU discovery fails.
The instance lifecycle already routes create/start/stop/delete through
CreateVGPU/DestroyVGPU, so dispatching vendor VFIO creates here would
activate the backend before assignment durability and release guards
exist. Reject vendor VFIO creates for now; destroy stays wired so
existing assignments remain releasable. The integration test skips on
vendor VFIO hosts at this layer and no longer asserts the transitional
stop-retention behavior.
Counting every free VF advertising a type overreports concurrent
capacity: sibling VFs share their parent GPU's framebuffer, so one 48Q
assignment revokes the type from every other VF on that GPU. Bound each
GPU's contribution by both its free VFs and how many times the profile
framebuffer fits into the GPU's remaining framebuffer, using the largest
still-creatable profile as a lower bound on what remains.
A single unreadable current_vgpu_type failed discoverVFs wholesale, and
GetGPUStatus turns a discovery error into a host with no GPU, so one
flaky sysfs read blanked out the host's entire GPU capacity for
admission and monitoring.
Skip unreadable VFs with a warning and keep the readable inventory: a
skipped VF is never selected for placement and never reconciled, both
safe directions. When no VF is readable, discovery still fails so a
wholesale sysfs outage cannot demote a vGPU host to passthrough while
assignments exist.
Also document that vendor VFIO vGPUs are known broken on Cloud
Hypervisor upstream and QEMU is the required hypervisor for GPU
instances.
listProfiles failed wholesale when one VF's creatable_vgpu_types read
failed, blanking every advertised profile while discoverVFs directly
above it already skips unreadable VFs for exactly that reason. Skip and
warn instead; underreporting is the safe direction for status and
admission.
Also document why openVFIOPaths stays strict where mdev's scan is lax
(it authorizes clearing a reused VF path), and the 0Q/0B parsing caveat
in framebufferFromProfileName.
listProfiles skips an unreadable VF but create still failed placement
wholesale when profileMetadata or selectLeastLoadedVF hit the same VF,
so /resources could advertise capacity a create then failed to use.
Skip the VF in both loops; it simply stops being a placement candidate.
A process that exits between the /proc listing and its fd walk surfaces
ENOENT or ESRCH; it holds nothing open, so skipping it cannot produce a
false "not in use" answer. Everything else still fails the scan closed.
An mdev assignment carrying neither a UUID nor a device path would
resolve to DestroyMdev("."); release nothing instead.
@yummybomb
yummybombforce-pushed the hypeship/vendor-vfio-backend branch from 0ddc917 to 6f98055CompareAugust 21, 2026 20:47
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@yummybomb@sjmiller609