Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

L2M: Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space &

L2M Logo

Welcome to the L2M repository! This is the official implementation of our ICCV'25 paper titled "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space".

Accepted to ICCV 2025 Conference

Quick Start

fromromatchimportl2mpp_modelimporttorchimportcv2device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
l2mpp=l2mpp_model(device=device, version="v2") # , version="v1"# Matchwarp, certainty=l2mpp.match("assets/sacre_coeur_A.jpg", "assets/sacre_coeur_B.jpg", device=device)
matches, certainty=l2mpp.sample(warp, certainty)
# Convert to pixel coordinates (RoMa produces matches in [-1,1]x[-1,1])kptsA, kptsB=roma_model.to_pixel_coordinates(matches, H_A, W_A, H_B, W_B)
# Find a fundamental matrix (or anything else of interest)F, mask=cv2.findFundamentalMat(
kptsA.cpu().numpy(), kptsB.cpu().numpy(), ransacReprojThreshold=0.2, method=cv2.USAC_MAGSAC, confidence=0.999999, maxIters=10000
)

🔗 Pretrained Model Weights

Pretrained checkpoints for L2M++ are publicly available on Hugging Face:

👉 Hugging Face:
https://huggingface.co/datasets/Liangyingping/L2Mpp-checkpoints

The model will automatically download the required weights when running inference.


🧠 Overview

Lift to Match (L2M) is a two-stage framework for dense feature matching that lifts 2D images into 3D space to enhance feature generalization and robustness. Unlike traditional methods that depend on multi-view image pairs, L2M is trained on large-scale, diverse single-view image collections.

🏗️ Data Generation

To enable training from single-view images, we simulate diverse multi-view observations and their corresponding dense correspondence labels in a fully automatic manner.

Stage 2.1: Novel View Synthesis

We lift a single-view image to a coarse 3D structure and then render novel views from different camera poses. These synthesized multi-view images are used to supervise the feature encoder with dense matching consistency.

Run the following to generate novel-view images with ground-truth dense correspondences:

python get_data.py \
--output_path [PATH-to-SAVE] \
--data_path [PATH-to-IMAGES] \
--disp_path [PATH-to-MONO-DEPTH]

This code provides an example on novel view generation with dense matching ground truth.

The disp_path should contain grayscale disparity maps predicted by Depth Anything V2 or another monocular depth estimator.

Below are examples of synthesized novel views with ground-truth dense correspondences, generated in Stage 2.1:


test_000002809

These demonstrate both the geometric diversity and high-quality pixel-level correspondence labels used for supervision.

For novel-view inpainting, we also provide a better inpainting model fine-tuned from Stable-Diffusion-2.0-Inpainting:

from diffusers import StableDiffusionInpaintPipeline
import torch
from diffusers.utils import load_image, make_image_grid
import PIL
model_path = "Liangyingping/Lift3Dreamer"
pipe = StableDiffusionInpaintPipeline.from_pretrained(
model_path, torch_dtype=torch.float16
)
pipe.to("cuda")
init_image = load_image("assets/debug_masked_image.png")
mask_image = load_image("assets/debug_mask.png")
W, H = init_image.size
prompt = "a photo of a person"
image = pipe(
prompt=prompt,
image=init_image,
mask_image=mask_image,
h=512, w=512
).images[0].resize((W, H))
print(image.size, init_image.size)
image2save = make_image_grid([init_image, mask_image, image], rows=1, cols=3)
image2save.save("image2save_ours.png")

If you use this inpainting model, please cite our paper:

@article{LIANG2026,
title = {Lift3Dreamer: Boosting Text-Driven Novel View Synthesis via Lifted 3D Inpainting Model from Single Images.},
journal = {Fundamental Research},
year = {2026},
issn = {2667-3258},
author = {Yingping Liang and Ying Fu and Jiaming Liu and Debing Zhang},
}

Stage 2.2: Relighting for Appearance Diversity

To improve feature robustness under varying lighting conditions, we apply a physics-inspired relighting pipeline to the synthesized 3D scenes.

Run the following to generate relit image pairs for training the decoder:

python relight.py

All outputs will be saved under the configured output directory, including original view, novel views, and their camera metrics with dense depth.

demo-data

Stage 2.3: Sky Masking

If desired, you can run sky_seg.py to mask out sky regions, which are typically textureless and not useful for matching. This can help reduce noise and focus training on geometrically meaningful regions.

python sky_seg.py

ADE_train_00000971

🙋‍♂️ Acknowledgements

We build upon recent advances in ROMA, GIM, MINIMA, and FiT3D.

Cite Our Paper

@inproceedings{Liang2025L2M,
author = {Yingping Liang and Yutao Hu and Wenqi Shao and Ying Fu},
title = {Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year = {2025},
pages = {6621--6631}
}

About

Official implementation of our ICCV'25 paper "Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space"

Resources

Stars

71 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages