Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision - #13

Merged
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master
Apr 27, 2022
Merged

Addition of mode "fastload", as well as support for 2-byte and 4-byte precision#13
leonbohmann merged 7 commits into
leonbohmann:devfrom
hakonbar:master

Conversation

@hakonbar

Copy link
Copy Markdown
Contributor

Kind of a big pull request, introducing two features at once. They've been implemented in such a way as to not interfere with existing functionality.

  • Added a 'fastload' mode, which takes advantage of the fact that consecutive data points in a measurement channel are stored as a contiguous "byte chunk" in the catman binary format instead of blockwise. You therefore only need to pass a pointer to the first byte as well as the length of the chunk.
  • Added the method "Channel.readExtHeader", in order to get at the attribute "ExportFormat". This attribute indicates the byte depth or precision of the measurement file, allowing the algorithm to differentiate.
  • Added the method "BinaryReader.read_float", which reads in 4-byte floating point numbers.
  • Changed the name of the method "read_single" to "read_byte" to avoid confusion with the newly added method.
  • Added some sample data from HBK with 2-, 4- and 8-byte data.

# Added a 'fastload' option, which reads the channel data to a numpy array.
# Added support for reading measurement data with 2-byte and 4-byte precision
- added the method 'read_float()'
- renamed the method 'read_single()' to 'read_byte()' to avoid confusion.
# Added test files with data at the different precision levels.
The code now sets the attribute "ExportFormat" to zero instead of throwing away the entire extended header.

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand correctly, "fastload" will instruct the Channel to load its data using Numpys Implementation of np.fromfile?

Do you think it would be a good idea to make fastload the default way of loading data? I mean if it does the same thing only faster... If so, you could change the default value for it.

Comment threadapread/entries.py

@leonbohmannleonbohmann left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok so all in all this seems fine. Only thing I don't get yet is the scale factor..

@leonbohmannleonbohmann added question Further information is requested enhancement New feature or request and removed question Further information is requested labels Apr 26, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

I am merging this into a new branch "dev" which I will use to develop the future release. Just to keep things organized!

@leonbohmann
leonbohmann changed the base branch from master to devApril 27, 2022 19:03
@leonbohmann
leonbohmann merged commit f37cc49 into leonbohmann:devApr 27, 2022
@leonbohmann

Copy link
Copy Markdown
Owner

The fastload option actually breaks the conversion into json format, because when using fastload it created the data objects of the channels as ndarray which is not marked as json-serializable.

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

Hi Leon, you could use a condition like "type(data) is ndarray" or "isinstance(data, ndarray)" to check if the data is stored in a numpy array, and if True, apply a statement like "data_to_write = data.tolist()" or "data_to_write = list(data)" before writing to json or csv.

For numerical data like we have here, however, it's usually better to use a binary format when writing to disk, as this requires a lot less storage space and makes for faster reading and writing of files. One suggestion here would be to implement a save method which uses the "pickle" module to pickle the whole object. This could be useful for users who want to use Python for further processing. One could also implement a method which writes the measurement data to parquet and the metadata to json. This would make the output more portable.

When working with homogeneous numerical data (where all values in the data structure are of the same datatype), numpy arrays are usually orders of magnitude faster than lists. The numpy and scipy libraries have also implemented heaps of functions which are optimized for just this data structure. You could therefore consider converting the measurement data to an ndarray also when not using fastload. I think your filtering function "lfilt()" might also return an ndarray, but I'm not sure.

@leonbohmann

Copy link
Copy Markdown
Owner

Yes you are right about that. Maybe I'll just overhaul the data datatype completely and switch to ndarray. Then it'll be upon the user to decide on how to save it.

This package should focus on reading the data only, most users will probably create theirnown plots and files eitherway...

@hakonbar

hakonbar commented May 20, 2022

Copy link
Copy Markdown
ContributorAuthor

In that case, the pickle module would be a perfect fit. It allows you to dump an item in your workspace to file with only a few lines of code. The file can then just as easily be loaded into the workspace again in a later Python session for further processing. See a code example below (excuse my Norwegian code):

`def lag_pickle(mappe_lagre,objekt,filnavn):

 if not filnavn.endswith('.pkl'):
from pathlib import Path
filnavn = Path(filnavn).stem + '.pkl'
with open(os.path.join(mappe_lagre,filnavn), 'wb') as outp:
pickle.dump(objekt, outp, pickle.HIGHEST_PROTOCOL)

def hent_pickle(mappe_last,filnavn):

 with open(os.path.join(mappe_last,filnavn), 'rb') as inp:
objekt = pickle.load(inp)
return objekt`

@hakonbar

Copy link
Copy Markdown
ContributorAuthor

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

@leonbohmann

Copy link
Copy Markdown
Owner

By the way, I've found a bug with the fastload mode which occurs when the file has fewer datapoints than is indicated in the header. The regular mode raises an IndexError there, but fastload doesn't, and produces gibberish instead. I'll try and fix it.

That will be a problem also for the reading using the original method. Therefor we should consider some error handljng to prevent the code failing using both methods..

@leonbohmann

leonbohmann commented May 20, 2022

Copy link
Copy Markdown
Owner

In that case, the pickle module would be a perfect fit....

Yes true. But I think this package should then only be used to convert the binary data to some ndarray in python. The seconds step will be up to the user.

While using the package myself, I realised that I tend to make a lot of changes in the package just so it fits my needs. It'll be more efficient, if we keep things and responsibilites simple, I think!

@leonbohmann

Copy link
Copy Markdown
Owner

New version is released containing your changes. I think the external header data is really helpful as well so I included that into the readme!

@LarissaPestana

Copy link
Copy Markdown

hello leon, is it possible to convert the read file into .xlxs?

@leonbohmann

Copy link
Copy Markdown
Owner

For questions and feature request please create a new issue.

Surely it is possible, but unfortunately that functionality is not part of this package. I did a quick search and found out, that you can convert a pandas dataframe to excel. For that, you would have to convert the channels to a dataframe first.

The other option would be to save the data as a csv file, you can simply open that with excel directly and save it as xlsx from there!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hakonbar@leonbohmann@LarissaPestana