Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

MatAnyone2Kit

code: GPL-3.0weights: S-Lab 1.0 non-commercial

A self-contained Swift package that runs the MatAnyone2 single-object video-matting model in real time on the Apple Neural Engine — stable 30 fps on an iPhone 16 (A18). The six precompiled Core ML models ship inside the package, so you add one dependency and feed it camera frames; there's no separate model download or launch-time compile.

Left: raw camera frame. Right: the same frame matted by MatAnyone2 on the Apple Neural Engine.

Same frame, split down the middle — raw camera on the left, MatAnyone2's real-time matte on the right. No green screen, no rotoscoping.

The conversion toolchain (and a writeup of every ANE/Swift optimization that made it real-time) lives in scripts/.

License: the package code is GPL-3.0, but the bundled MatAnyone2 weights are NTU S-Lab License 1.0 — non-commercial only. Using this package with the bundled weights is non-commercial. Details ↓

Install

Swift Package Manager — add it as a dependency in your Package.swift:

.package(url:"https://github.com/flowtyone/MatAnyone2Kit", from:"1.0.0")
// then add "MatAnyoneKitCoreML" to your target's dependencies

Or in Xcode: File ▸ Add Package Dependencies… and paste https://github.com/flowtyone/MatAnyone2Kit.

Requires iOS 17+ / macOS 14+ (the on-device ANE-eligibility probe uses MLComputePlan, iOS 17.4 / macOS 14.4, guarded internally).

Use

import MatAnyoneKitCoreML
// Loads the bundled, precompiled models; nil if they can't be loaded.
guardlet matte =MatAnyoneMatte()else{return}
// On each camera frame (BGRA CVPixelBuffer). The first frame with a person seeds the
// tracker via Vision; subsequent frames track from MatAnyone's own memory.
matte.matte(pixelBuffer){ frame inguardlet alpha = frame.alpha else{
// .passthrough — no matte yet (no person seeded). Show the full camera frame.
return}
// `alpha` is a single-channel CVPixelBuffer (OneComponent8, 1 = foreground) at the model
// working resolution (288×512), framed like the camera frame. Composite it however you like.
iflet timing = frame.timing {
// per-stage wall-clock (ms): timing.preprocessMs / inferenceMs / postprocessMs
}}

matte(_:completion:) runs synchronously on the calling queue (typically your camera/video output queue) and calls completion once. For best throughput, drive it back-to-back keeping only the freshest pending frame rather than once per camera tick.

Real-time best practices

These are the settings that get a stable 30 fps on an iPhone 16 (A18) in a live camera app.

Resolution. The bundled models run at a fixed 288×512 (portrait) — that's the working size tuned for 30 fps on the ANE, and you don't choose it (matte.workingWidth/workingHeight report it). So you control cost upstream, at capture: feed 720p or lower. Capturing at 4K just burns ISP and conversion time for frames the matte immediately downsizes to 288×512. Capture BGRA (kCVPixelFormatType_32BGRA), upright/portrait — the returned alpha is framed identically to the frame you pass in, so you composite it with the same aspect-fill UVs as the camera.

session.sessionPreset =.hd1280x720 // 720p is plenty; the matte works at 288×512
output.videoSettings =[kCVPixelBufferPixelFormatTypeKey asString: kCVPixelFormatType_32BGRA]
output.alwaysDiscardsLateVideoFrames =true

Capture at 60 fps, not 30. The matte's throughput quantum is the camera tick. At 30 fps (33 ms ticks) a 34 ms matte misses every tick and snaps to 15 fps. Pinning the camera to 60 fps (16.7 ms ticks) lets a sub-33 ms matte hold 30 fps, with a 20 fps floor otherwise. Pinning min == max also stops iOS from auto-dropping the rate in low light.

iflet maxRate = device.activeFormat.videoSupportedFrameRateRanges.map(\.maxFrameRate).max(),(try? device.lockForConfiguration())!=nil{letdur=CMTime(value:1, timescale:CMTimeScale(min(60.0, maxRate).rounded()))
device.activeVideoMinFrameDuration = dur
device.activeVideoMaxFrameDuration = dur
device.unlockForConfiguration()}

Drive it back-to-back, don't drop on the camera tick. Run the matte off the main thread on one serial queue, and when a frame arrives mid-pass keep only the freshest one, then process it the instant the current pass finishes. This decouples throughput from the 33 ms delivery window (a 35 ms matte streams at ~28 fps instead of collapsing to 15) while keeping latency at ~one frame:

finalclassMattePacer{privateletmatte:MatAnyoneMatteprivateletqueue=DispatchQueue(label:"matte", qos:.userInitiated)privateletlock=NSLock()privatevarbusy=falseprivatevarpending:CVPixelBuffer?varonResult:((CVPixelBuffer,MatAnyoneMatte.MatteFrame)->Void)?init(_ matte:MatAnyoneMatte){self.matte = matte }
// Call from your camera delegate on every frame.
func submit(_ frame:CVPixelBuffer){
lock.lock()if busy { pending = frame; lock.unlock(); return} // keep only the freshest
busy =true; lock.unlock()process(frame)}privatefunc process(_ frame:CVPixelBuffer){
queue.async{[self]in
matte.matte(frame){ result inonResult?(frame, result) // source frame + its matte, in sync
lock.lock(); letnext= pending; pending =nilif next ==nil{ busy =false}; lock.unlock()iflet next {process(next)} // run now, not on the next camera tick
}}}}

Composite the frame with its own matte.onResult hands you the exact source frame each matte was computed from — composite that pair so the cutout never lags the camera by a frame.

Load off the main thread. First-launch ANE specialization takes a few seconds; construct MatAnyoneMatte() on a background task and run a passthrough (full-frame camera) until it's ready, so your preview appears instantly.

Task.detached(priority:.utility){letmatte=MatAnyoneMatte(); /* swap it in when ready */ }

Mind the thermal envelope. The matte runs on the ANE, leaving the GPU free — but a heavy GPU renderer plus the matte plus the camera ISP sustained together will throttle the device before ProcessInfo.thermalState even reports fair. Watch per-stage latency, not just thermalState.

What's inside

typerole
MatAnyoneMattetop-level facade: Vision seeding + stateful tracking, returns MatteFrame
MatAnyoneCoreMLEnginestateful inference loop over the six models
MatAnyoneCoreMLmodel loading + efficient MLMultiArray[Float] conversion
MemoryBank / MemoryMathkey/affinity memory, top-k softmax readout (Accelerate, parallelized)

The lower-level engine (MatAnyoneCoreMLEngine, MatAnyoneCoreML) is public if you want to drive the pipeline directly; most callers only need MatAnyoneMatte.

Configuration

MatAnyoneMatte.diagnostics =true // verbose per-frame logging + ANE-eligibility probe
MatAnyoneMatte.defaultUnit =.cpuAndNeuralEngine
MatAnyoneMatte.unitOverrides =["objsummary":.cpuAndGPU] // per-model compute placement
// Load models from a custom directory instead of the bundled ones:
letmatte=MatAnyoneMatte(modelsDir: someURL)

Set these before constructing MatAnyoneMatte.

License

MatAnyone2Kit ships under two licenses — see NOTICE.md for the full breakdown:

  • Swift package source code — GNU GPL-3.0 (LICENSE).
  • Bundled MatAnyone2 model weights (the Core ML models in Sources/MatAnyoneKitCoreML/Resources/MatAnyone/) — NTU S-Lab License 1.0, non-commercial only. These are a Core ML conversion of the MatAnyone2 weights by S-Lab, NTU; converting them does not change their license.

⚠️Using this package with the bundled weights is non-commercial only. The GPL-3.0 on the code does not grant any commercial rights to the weights. For commercial use of the weights, contact the authors (see the weights NOTICE.md).

About

No description, website, or topics provided.

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages