Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use MirroredArray instead of numpy arrays to reduce boilerplate - #671

Closed
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array
Closed

Use MirroredArray instead of numpy arrays to reduce boilerplate#671
connorjward wants to merge 14 commits into
gpufrom
gpu-clever-array

Conversation

@connorjward

Copy link
Copy Markdown
Collaborator

@kaushikcfd

@JDBetteridge and I took a harder look at your code yesterday and came up with an idea that we think could reduce a large amount of boilerplate.

The key idea is the introduction of what I have called a MirroredArray. It is effectively a host/device 'aware' version of a numpy array. It would have a few responsibilities:

  • Ensure that the data on the host/device is up-to-date when needed
  • Do the correct transformation to a PETSc Vec depending on the backend
  • Return an appropriate pointer for passing to the kernel (_kernel_args_), depending on whether or not offloading is enabled

I think this would be good for removing lots of boilerplate because it would (I think) remove the need for us to have separate implementations of Dat, ExtrudedSet, Global, etc per backend. All that would need to be done instead is to replace any numpy arrays that we want to exist on both host and device with MirroredArrays.

What I've provided here is quite a rough sketch of what I think such a solution would look like. The key thing I have not yet tackled is how we might transform these arrays into PETSc Vecs with context managers and such.

Please let me know your thoughts. @JDBetteridge and I would be very happy to have a call with you at some point to discuss it further.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was just a hack to create a global queue object since I think a compute_backend object may not be required.

Comment threadpyop2/array.py
self.device_to_host_copy()

@property
def vec(self):

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Subclasses would need to implement the right thing here. I have not tried to make my code work for can_be_represented_as_petscvec but I don't think it would require a radical rethink.

Comment threadpyop2/types/map.py
(iterset.total_size, arity), allow_none=True)
self.shape = (iterset.total_size, arity)
shape = (iterset.total_size, arity)
self._values_array = MirroredArray.new(values, dtypes.IntType, shape)

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note how if we do this then backends.opencl.Map can go away.

@kaushikcfdkaushikcfd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Some of the logic here would force unnecessary host<->device copies, making it hard to evaluate purely from the decrease in the LOC. I agree that there is some duplication between the OpenCL and CUDA backends. But not sure whether this MirroredArray abstraction is the way to go.

(Yep we should schedule a call!)

Comment threadpyop2/array.py
self._host_data = np.zeros(shape, dtype=dtype)
else:
self._host_data = verify_reshape(data, dtype, shape)
self.availability = ON_BOTH

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the availability ON_BOTH here?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is probably wrong if data is not None. For simplicity I was assuming that the array would be initialised with a valid copy on both host and device. This definitely doesn't need to be the case (and I doubt I have implemented it correctly anyway).

Comment threadpyop2/array.py
Comment on lines +48 to +49
# lazy but for now assume that the data is always modified if we access
# the pointer

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree we should do this right now. But probably we could patch PyOP2 in the future where a ParLoop tells us what is the access type.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the parloop does tell you the access type? otherwise nothing would work

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A simple solution here would be to replace array.kernel_arg with array.get_kernel_arg(access) which would return either array.{host,device}_ptr_ro or array.{host,device}_ptr as appropriate. I've actually commented this out below.

Comment threadpyop2/array.py
Comment on lines +136 to +145
if configuration["backend"] == "OPENCL":
# TODO: Instruct the user to pass
# -viennacl_backend opencl
# -viennacl_opencl_device_type gpu
# create a dummy vector and extract its associated command queue
x = PETSc.Vec().create(PETSc.COMM_WORLD)
x.setType("viennacl")
x.setSizes(size=1)
queue_ptr = x.getCLQueueHandle()
cl_queue = pyopencl.CommandQueue.from_int_ptr(queue_ptr, retain=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm assuming this would go away.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly. What I'm trying to illustrate here is that we can do without the OpenCLBackend class since we don't need to maintain different Dat, Global, etc subclasses.

Comment threadpyop2/array.py
Comment on lines +94 to +98
self.ensure_availability_on_host()
self.availability = ON_HOST
v = self._host_data.view()
v.setflags(write=True)
return v

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This would force a device->host, which isn't necessary and would be performance limiting.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a device -> host copy is needed here as otherwise the numpy array you get back might be wrong?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we should return the array on the device (cl.array.Array, pycuda.GPUArray) whenever we are in the offloading context. They allow us to perform numpy-like operations on the device, IMO no good reason for returning the numpy array. Wdyt?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That makes a lot of sense. I was assuming here that every time we called this function we would want to inspect the data, which would require a copy to the device.

Comment threadpyop2/array.py
Comment on lines +151 to +154
if data is None:
self._device_data = pyopencl.array.empty(cl_queue, shape, dtype)
else:
self._device_data = pyopencl.array.to_device(cl_queue, data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Minor]: This is probably missing a super().__init__(data, shape, dtype).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. Good spot.

Comment threadpyop2/backends/cpu.py

if mpi.MPI.VERSION >= 3:
requests.append(self.comm.Iallreduce(glob._data,
requests.append(self.comm.Iallreduce(glob.data,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not convinced this is correct. glob.data shouldn't necessarily return an array on the host.

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought that MPI routines occur between host copies. Is that not the case?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might want to resolve #671 (comment) first as we are going back and forth on what Global.data should actually return.

Comment threadpyop2/array.py
Comment on lines +167 to +171
def host_to_device_copy(self):
self._device_data.set(self._host_data)

def device_to_host_copy(self):
self._device_data.get(self._host_data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The logic here would be different based on where this is a Dat or (Map|Global). For Dats we need to handle the special case when we don't want to synchronize the halo values as pointed by Lawrence in #574 (comment).

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose. Perhaps a MirroredArrayWithHalo would be the way to go.

@kaushikcfd
kaushikcfdforce-pushed the gpu branch 3 times, most recently from 5bed614 to 3df2554CompareNovember 18, 2022 00:16
@connorjward

Copy link
Copy Markdown
CollaboratorAuthor

Closing as superseded by #691.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@connorjward@wence-@kaushikcfd@JDBetteridge