Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

141 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

GetPy - A Vectorized Python Dict/Set

The goal of GetPy is to provide the highest performance python dict/set that integrates into the python scientific ecosystem.

Installation

pip install getpy

Note only a linux build is currently distributed. If you would like to build the package from source you can clone the repo and run python setup.py install. Compilation will require 16gb of ram. I am working on getting that down.

About

GetPy is a thin binding to the Parallel Hashmap (https://github.com/greg7mdp/parallel-hashmap.git) which is the current state of the art unordered map/set with minimal memory overhead and fast runtime speed. The binding layer is supported by PyBind11 (https://github.com/pybind/pybind11.git) which is fast to compile and simple to extend.

How To Use

The gp.Dict and gp.Set objects are designed to maintain a similar interface to the corresponding standard python objects. There are some key differences though, which are necessary for vectorization and other performance considerations.

  1. gp.Dict.__init__ has three arguments key_type, value_type, and default_value. The type arguments are define which compiled data structure will be used under the hood, and the full list of preset combinations of np.dtypes is found with gp.dict_types. You can also specify a default_value at construction which must be castable to the value_type. This is the value returned by the dictionary if a key is not found.

  2. All of getpy.Dict methods support a vectorized interface. Therefore, methods like gp.Dict.__getitem__, gp.Dict.__setitem__, and gp.Dict.__delitem__ can be performed with an np.ndarray. That allows the performance critical for-loop to happen within the compiled c++. Note that some dunder methods cannot be vectorized such as __contains__. Therefore, some keywords like in do not behave as expected. Those methods are renamed without the double underscores to note their deviation from the standard interface.

  3. If a key does not exist, gp.Dict.__getitem__ will return the default_value. If you do not specify the default_value, it will default to the default constructor of your data type (all 0 bits). If you would like to know the difference between a key that does not exist and a key that returns the default value, you should first run gp.contains on your key/array of keys, and then retrieve values corresponding to keys that exist.

  4. There is also a gp.MultiDict object. This object stores multiple unique values per key.

Examples

Simple Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Default Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type)
gp_dict=gp.Dict(key_type, value_type, default_value=42)
gp_dict[keys] =valuesrandom_keys=np.random.randint(1, 1000, size=500, dtype=key_type)
random_values=gp_dict[random_keys]

Byteset Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('S8')
value_type=np.dtype('S8')
keys=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=key_type)
values=np.array([np.random.bytes(8) foriinrange(10**2)], dtype=value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Multidimensional Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=key_type).reshape(10,10)
values=np.random.randint(1, 1000, size=10**2, dtype=value_type).reshape(10,10)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =values

Bitpack Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**2, dtype=np.dtype('u2')).reshape(25,4).view(key_type)
values=np.random.randint(1, 1000, size=(10**2)/2, dtype=np.dtype('u4')).reshape(25,2).view(value_type)
gp_dict=gp.Dict(key_type, value_type)
gp_dict[keys] =valuesunpacked_values=gp_dict[keys].view(np.dtype('u4'))

Serialization Example

importnumpyasnpimportgetpyasgpkey_type=np.dtype('u8')
value_type=np.dtype('u8')
keys=np.random.randint(1, 1000, size=10**1, dtype=key_type)
values=np.random.randint(1, 1000, size=10**1, dtype=value_type)
gp_dict_1=gp.Dict(key_type, value_type)
gp_dict_1[keys] =valuesgp_dict_1.dump('test/test.hashtable.bin')
gp_dict_2=gp.Dict(key_type, value_type)
gp_dict_2.load('test/test.hashtable.bin')

Supported Data Types

dict_types= {
(np.dtype('u4'), np.dtype('u1')) : _gp.Dict_u4_u1,
(np.dtype('u4'), np.dtype('u2')) : _gp.Dict_u4_u2,
(np.dtype('u4'), np.dtype('u4')) : _gp.Dict_u4_u4,
(np.dtype('u4'), np.dtype('u8')) : _gp.Dict_u4_u8,
(np.dtype('u4'), np.dtype('i1')) : _gp.Dict_u4_i1,
(np.dtype('u4'), np.dtype('i2')) : _gp.Dict_u4_i2,
(np.dtype('u4'), np.dtype('i4')) : _gp.Dict_u4_i4,
(np.dtype('u4'), np.dtype('i8')) : _gp.Dict_u4_i8,
(np.dtype('u4'), np.dtype('f4')) : _gp.Dict_u4_f4,
(np.dtype('u4'), np.dtype('f8')) : _gp.Dict_u4_f8,
(np.dtype('u4'), np.dtype('S8')) : _gp.Dict_u4_S8,
(np.dtype('u4'), np.dtype('S16')) : _gp.Dict_u4_S16,
(np.dtype('u8'), np.dtype('u1')) : _gp.Dict_u8_u1,
(np.dtype('u8'), np.dtype('u2')) : _gp.Dict_u8_u2,
(np.dtype('u8'), np.dtype('u4')) : _gp.Dict_u8_u4,
(np.dtype('u8'), np.dtype('u8')) : _gp.Dict_u8_u8,
(np.dtype('u8'), np.dtype('i1')) : _gp.Dict_u8_i1,
(np.dtype('u8'), np.dtype('i2')) : _gp.Dict_u8_i2,
(np.dtype('u8'), np.dtype('i4')) : _gp.Dict_u8_i4,
(np.dtype('u8'), np.dtype('i8')) : _gp.Dict_u8_i8,
(np.dtype('u8'), np.dtype('f4')) : _gp.Dict_u8_f4,
(np.dtype('u8'), np.dtype('f8')) : _gp.Dict_u8_f8,
(np.dtype('u8'), np.dtype('S8')) : _gp.Dict_u8_S8,
(np.dtype('u8'), np.dtype('S16')) : _gp.Dict_u8_S16,
(np.dtype('i4'), np.dtype('u1')) : _gp.Dict_i4_u1,
(np.dtype('i4'), np.dtype('u2')) : _gp.Dict_i4_u2,
(np.dtype('i4'), np.dtype('u4')) : _gp.Dict_i4_u4,
(np.dtype('i4'), np.dtype('u8')) : _gp.Dict_i4_u8,
(np.dtype('i4'), np.dtype('i1')) : _gp.Dict_i4_i1,
(np.dtype('i4'), np.dtype('i2')) : _gp.Dict_i4_i2,
(np.dtype('i4'), np.dtype('i4')) : _gp.Dict_i4_i4,
(np.dtype('i4'), np.dtype('i8')) : _gp.Dict_i4_i8,
(np.dtype('i4'), np.dtype('f4')) : _gp.Dict_i4_f4,
(np.dtype('i4'), np.dtype('f8')) : _gp.Dict_i4_f8,
(np.dtype('i4'), np.dtype('S8')) : _gp.Dict_i4_S8,
(np.dtype('i4'), np.dtype('S16')) : _gp.Dict_i4_S16,
(np.dtype('i8'), np.dtype('u1')) : _gp.Dict_i8_u1,
(np.dtype('i8'), np.dtype('u2')) : _gp.Dict_i8_u2,
(np.dtype('i8'), np.dtype('u4')) : _gp.Dict_i8_u4,
(np.dtype('i8'), np.dtype('u8')) : _gp.Dict_i8_u8,
(np.dtype('i8'), np.dtype('i1')) : _gp.Dict_i8_i1,
(np.dtype('i8'), np.dtype('i2')) : _gp.Dict_i8_i2,
(np.dtype('i8'), np.dtype('i4')) : _gp.Dict_i8_i4,
(np.dtype('i8'), np.dtype('i8')) : _gp.Dict_i8_i8,
(np.dtype('i8'), np.dtype('f4')) : _gp.Dict_i8_f4,
(np.dtype('i8'), np.dtype('f8')) : _gp.Dict_i8_f8,
(np.dtype('i8'), np.dtype('S8')) : _gp.Dict_i8_S8,
(np.dtype('i8'), np.dtype('S16')) : _gp.Dict_i8_S16,
(np.dtype('S8'), np.dtype('u1')) : _gp.Dict_S8_u1,
(np.dtype('S8'), np.dtype('u2')) : _gp.Dict_S8_u2,
(np.dtype('S8'), np.dtype('u4')) : _gp.Dict_S8_u4,
(np.dtype('S8'), np.dtype('u8')) : _gp.Dict_S8_u8,
(np.dtype('S8'), np.dtype('i1')) : _gp.Dict_S8_i1,
(np.dtype('S8'), np.dtype('i2')) : _gp.Dict_S8_i2,
(np.dtype('S8'), np.dtype('i4')) : _gp.Dict_S8_i4,
(np.dtype('S8'), np.dtype('i8')) : _gp.Dict_S8_i8,
(np.dtype('S8'), np.dtype('f4')) : _gp.Dict_S8_f4,
(np.dtype('S8'), np.dtype('f8')) : _gp.Dict_S8_f8,
(np.dtype('S8'), np.dtype('S8')) : _gp.Dict_S8_S8,
(np.dtype('S8'), np.dtype('S16')) : _gp.Dict_S8_S16,
(np.dtype('S16'), np.dtype('u1')) : _gp.Dict_S16_u1,
(np.dtype('S16'), np.dtype('u2')) : _gp.Dict_S16_u2,
(np.dtype('S16'), np.dtype('u4')) : _gp.Dict_S16_u4,
(np.dtype('S16'), np.dtype('u8')) : _gp.Dict_S16_u8,
(np.dtype('S16'), np.dtype('i1')) : _gp.Dict_S16_i1,
(np.dtype('S16'), np.dtype('i2')) : _gp.Dict_S16_i2,
(np.dtype('S16'), np.dtype('i4')) : _gp.Dict_S16_i4,
(np.dtype('S16'), np.dtype('i8')) : _gp.Dict_S16_i8,
(np.dtype('S16'), np.dtype('f4')) : _gp.Dict_S16_f4,
(np.dtype('S16'), np.dtype('f8')) : _gp.Dict_S16_f8,
(np.dtype('S16'), np.dtype('S8')) : _gp.Dict_S16_S8,
(np.dtype('S16'), np.dtype('S16')) : _gp.Dict_S16_S16,
}
set_types= {
np.dtype('u4') : _gp.Set_u4,
np.dtype('u8') : _gp.Set_u8,
np.dtype('i4') : _gp.Set_i4,
np.dtype('i8') : _gp.Set_i8,
np.dtype('S8') : _gp.Set_S8,
np.dtype('S16') : _gp.Set_S16,
}

About

A Vectorized Python Dict/Set

Resources

Stars

115 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages