Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo.png?raw=true#gh-light-mode-only

https://github.com/unifyai/unifyai.github.io/blob/main/img/externally_linked/logo_dark.png?raw=true#gh-dark-mode-only



3D Vision functions with end-to-end support for machine learning developers, written in Ivy.

Contents

Overview

What is Ivy Vision?

Ivy vision focuses predominantly on 3D vision, with functions for camera geometry, image projections, co-ordinate frame transformations, forward warping, inverse warping, optical flow, depth triangulation, voxel grids, point clouds, signed distance functions, and others. Check out the docs for more info!

The library is built on top of the Ivy machine learning framework. This means all functions simultaneously support: Jax, Tensorflow, PyTorch, MXNet, and Numpy.

Ivy Libraries

There are a host of derived libraries written in Ivy, in the areas of mechanics, 3D vision, robotics, gym environments, neural memory, pre-trained models + implementations, and builder tools with trainers, data loaders and more. Click on the icons below to learn more!








Quick Start

Ivy vision can be installed like so: pip install ivy-vision==0.0.1.post0

To quickly see the different aspects of the library, we suggest you check out the demos! we suggest you start by running the script run_through.py, and read the "Run Through" section below which explains this script.

For more interactive demos, we suggest you run either coords_to_voxel_grid.py or render_image.py in the interactive demos folder.

Run Through

We run through some of the different parts of the library via a simple ongoing example script. The full script is available in the demos folder, as file run_through.py. First, we select a random backend framework to use for the examples, from the options ivy.jax, ivy.tensorflow, ivy.torch, ivy.mxnet or ivy.numpy, and use this to set the ivy backend framework.

importivyivy.set_backend(ivy.choose_random_backend())

Camera Geometry

To get to grips with some of the basics, we next show how to construct ivy containers which represent camera geometry. The camera intrinsic matrix, extrinsic matrix, full matrix, and all of their inverses are central to most of the functions in this library.

All of these matrices are contained within the Ivy camera geometry class.

# intrinsics# common intrinsic paramsimg_dims= [512, 512]
pp_offsets=ivy.array([dim/2-0.5fordiminimg_dims], 'float32')
cam_persp_angles=ivy.array([60*np.pi/180] *2, 'float32')
# ivy cam intrinsics containerintrinsics=ivy_vision.persp_angles_and_pp_offsets_to_intrinsics_object(
cam_persp_angles, pp_offsets, img_dims)
# extrinsics# 3 x 4cam1_inv_ext_mat=ivy.array(np.load(data_dir+'/cam1_inv_ext_mat.npy'), 'float32')
cam2_inv_ext_mat=ivy.array(np.load(data_dir+'/cam2_inv_ext_mat.npy'), 'float32')
# full geometry# ivy cam geometry containercam1_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam1_inv_ext_mat, intrinsics)
cam2_geom=ivy_vision.inv_ext_mat_and_intrinsics_to_cam_geometry_object(
cam2_inv_ext_mat, intrinsics)
cam_geoms= [cam1_geom, cam2_geom]

The geometries used in this quick start demo are based upon the scene presented below.

https://github.com/unifyai/vision/blob/main/docs/images/scene.png?raw=true

The code sample below demonstrates all of the attributes contained within the Ivy camera geometry class.

forcam_geomincam_geoms:
assertcam_geom.intrinsics.focal_lengths.shape== (2,)
assertcam_geom.intrinsics.persp_angles.shape== (2,)
assertcam_geom.intrinsics.pp_offsets.shape== (2,)
assertcam_geom.intrinsics.calib_mats.shape== (3, 3)
assertcam_geom.intrinsics.inv_calib_mats.shape== (3, 3)
assertcam_geom.extrinsics.cam_centers.shape== (3, 1)
assertcam_geom.extrinsics.Rs.shape== (3, 3)
assertcam_geom.extrinsics.inv_Rs.shape== (3, 3)
assertcam_geom.extrinsics.ext_mats_homo.shape== (4, 4)
assertcam_geom.extrinsics.inv_ext_mats_homo.shape== (4, 4)
assertcam_geom.full_mats_homo.shape== (4, 4)
assertcam_geom.inv_full_mats_homo.shape== (4, 4)

Load Images

We next load the color and depth images corresponding to the two camera frames. We also construct the depth-scaled homogeneous pixel co-ordinates for each image, which is a central representation for the ivy_vision functions. This representation simplifies projections between frames.

# load images# h x w x 3color1=ivy.array(cv2.imread(data_dir+'/rgb1.png').astype(np.float32) /255)
color2=ivy.array(cv2.imread(data_dir+'/rgb2.png').astype(np.float32) /255)
# h x w x 1depth1=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth1.png', -1).tobytes(), np.float32), img_dims+ [1]))
depth2=ivy.array(np.reshape(np.frombuffer(cv2.imread(
data_dir+'/depth2.png', -1).tobytes(), np.float32), img_dims+ [1]))
# depth scaled pixel coords# h x w x 3u_pix_coords=ivy_vision.create_uniform_pixel_coords_image(img_dims)
ds_pixel_coords1=u_pix_coords*depth1ds_pixel_coords2=u_pix_coords*depth2

The rgb and depth images are presented below.

https://github.com/unifyai/vision/blob/main/docs/images/rgb_and_depth.png?raw=true

Optical Flow and Depth Triangulation

Now that we have two cameras, their geometries, and their images fully defined, we can start to apply some of the more interesting vision functions. We start with some optical flow and depth triangulation functions.

# required mat formatscam1to2_full_mat_homo=ivy.matmul(cam2_geom.full_mats_homo, cam1_geom.inv_full_mats_homo)
cam1to2_full_mat=cam1to2_full_mat_homo[..., 0:3, :]
full_mats_homo=ivy.concat((ivy.expand_dims(cam1_geom.full_mats_homo, axis=0),
ivy.expand_dims(cam2_geom.full_mats_homo, axis=0)), axis=0)
full_mats=full_mats_homo[..., 0:3, :]
# flowflow1to2=ivy_vision.flow_from_depth_and_cam_mats(ds_pixel_coords1, cam1to2_full_mat)
# depth againdepth1_from_flow=ivy_vision.depth_from_flow_and_cam_mats(flow1to2, full_mats)

Visualizations of these images are given below.

https://github.com/unifyai/vision/blob/main/docs/images/flow_and_depth.png?raw=true

Inverse and Forward Warping

Most of the vision functions, including the flow and depth functions above, make use of image projections, whereby an image of depth-scaled homogeneous pixel-coordinates is transformed into cartesian co-ordinates relative to the acquiring camera, the world, another camera, or transformed directly to pixel co-ordinates in another camera frame. These projections also allow warping of the color values from one camera to another.

For inverse warping, we assume depth to be known for the target frame. We can then determine the pixel projections into the source frame, and bilinearly interpolate these color values at the pixel projections, to infer the color image in the target frame.

Treating frame 1 as our target frame, we can use the previously calculated optical flow from frame 1 to 2, in order to inverse warp the color data from frame 2 to frame 1, as shown below.

# inverse warp renderingwarp=u_pix_coords[..., 0:2] +flow1to2color2_warp_to_f1=ivy_vision.image.bilinear_resample(color2, warp)
# projected depth scaled pixel coords 2ds_pixel_coords1_wrt_f2=ivy_vision.ds_pixel_to_ds_pixel_coords(ds_pixel_coords1, cam1to2_full_mat)
# projected depth 2depth1_wrt_f2=ds_pixel_coords1_wrt_f2[..., -1:]
# inverse warp depthdepth2_warp_to_f1=ivy_vision.image.bilinear_resample(depth2, warp)
# depth validitydepth_validity=ivy.abs(depth1_wrt_f2-depth2_warp_to_f1) <0.01# inverse warp rendering with maskcolor2_warp_to_f1_masked=ivy.where(depth_validity, color2_warp_to_f1, ivy.zeros_like(color2_warp_to_f1))

Again, visualizations of these images are given below. The images represent intermediate steps for the inverse warping of color from frame 2 to frame 1, which is shown in the bottom right corner.

https://github.com/unifyai/vision/blob/main/docs/images/inverse_warped.png?raw=true

For forward warping, we instead assume depth to be known in the source frame. A common approach is to construct a mesh, and then perform rasterization of the mesh.

The ivy method ivy_vision.render_pixel_coords instead takes a simpler approach, by determining the pixel projections into the target frame, quantizing these to integer pixel co-ordinates, and scattering the corresponding color values directly into these integer pixel co-ordinates.

This process in general leads to holes and duplicates in the resultant image, but when compared to inverse warping, it has the beneft that the target frame does not need to correspond to a real camera with known depth. Only the target camera geometry is required, which can be for any hypothetical camera.

We now consider the case of forward warping the color data from camera frame 2 to camera frame 1, and again render the new color image in target frame 1.

# forward warp renderingds_pixel_coords1_proj=ivy_vision.ds_pixel_to_ds_pixel_coords(
ds_pixel_coords2, ivy.inv(cam1to2_full_mat_homo)[..., 0:3, :])
depth1_proj=ds_pixel_coords1_proj[..., -1:]
ds_pixel_coords1_proj=ds_pixel_coords1_proj[..., 0:2] /depth1_projfeatures_to_render=ivy.concat((depth1_proj, color2), axis=-1)
# without depth bufferf1_forward_warp_no_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=False)
# with depth bufferf1_forward_warp_w_db, _, _=ivy_vision.quantize_to_image(
ivy.reshape(ds_pixel_coords1_proj, (-1, 2)), img_dims, ivy.reshape(features_to_render, (-1, 4)),
ivy.zeros_like(features_to_render), with_db=Falseifivy.get_framework() =='mxnet'elseTrue)

Again, visualizations of these images are given below. The images show the forward warping of both depth and color from frame 2 to frame 1, which are shown with and without depth buffers in the right-hand and central columns respectively.

https://github.com/unifyai/vision/blob/main/docs/images/forward_warped.png?raw=true

Interactive Demos

In addition to the examples above, we provide two further demo scripts, which are more visual and interactive, and are each built around a particular function.

Rather than presenting the code here, we show visualizations of the demos. The scripts for these demos can be found in the interactive demos folder.

Neural Rendering

The first demo uses method ivy_vision.render_implicit_features_and_depth to train a Neural Radiance Field (NeRF) model to encode a lego digger. The NeRF model can then be queried at new camera poses to render new images from poses unseen during training.

Co-ordinates to Voxel Grid

The second demo captures depth and color images from a set of cameras, converts the depth to world-centric co-ordinartes, and uses the method ivy_vision.coords_to_voxel_grid to voxelize the depth and color values into a grid, as shown below:

Point Rendering

The final demo again captures depth and color images from a set of cameras, but this time uses the method ivy_vision.quantize_to_image to dynamically forward warp and point render the images into a new target frame, as shown below. The acquiring cameras all remain static, while the target frame for point rendering moves freely.

Get Involved

We hope the functions in this library are useful to a wide range of machine learning developers. However, there are many more areas of 3D vision which could be covered by this library.

If there are any particular vision functions you feel are missing, and your needs are not met by the functions currently on offer, then we are very happy to accept pull requests!

We look forward to working with the community on expanding and improving the Ivy vision library.

Citation

@article{lenton2021ivy,
title={Ivy: Templated deep learning for inter-framework portability},
author={Lenton, Daniel and Pardo, Fabio and Falck, Fabian and James, Stephen and Clark, Ronald},
journal={arXiv preprint arXiv:2102.02886},
year={2021}
}

About

3D Vision functions with end-to-end support for deep learning developers, written in Ivy.

Topics

Resources

Stars

72 stars

Watchers

14 watching

Forks

Releases

Packages

Used by

Contributors

Languages