Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

LZO for Java

Introduction

There is no version of LZO in pure Java. The obvious solution is to take the C source code, and feed it to the Java compiler, modifying the Java compiler as necessary to make it compile.

This package is an implementation of that obvious solution, for which I can only apologise to the world.

It turns out, however, that the compression performance on a single 2.4GHz laptop CPU is in excess of 500Mb/sec, and decompression runs at 815Mb/sec, which seems to be more than adequate. Run PerformanceTest on an appropriate file to reproduce these figures.

Example

Compression:

	OutputStream out = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoCompressor compressor = LzoLibrary.getInstance().newCompressor(algorithm, null);
LzoOutputStream stream = new LzoOutputStream(out, compressor, 256);
stream.write(...);

Decompression:

	InputStream in = ...;
LzoAlgorithm algorithm = LzoAlgorithm.LZO1X;
LzoDecompressor decompressor = LzoLibrary.getInstance().newDecompressor(algorithm, null);
LzoInputStream stream = new LzoInputStream(in, decompressor);
stream.read(...);

Documentation

The JavaDoc API is available.

Hadoop Notes

Notes on BlockCompressionStream, as of Hadoop 0.21.x:

  • If you write 1 byte, then a large block, BlockCompressorStream will flush the single-byte block before compressing the large block. This is inefficient.

  • If you write a large block to a fresh stream, BlockCompressorStream will flush existing data, which will write a zero uncompressed length to the file, but follow it with no blocks, thus breaking the ulen-clen-data format. This is wrong. There is no contract for the finished() method to avoid this, since it must return false at the top of write(), then must (with no other mutator calls) return true in BlockCompressorStream.finish() in order to avoid the empty block; having returned true there, compress() must be able to return a nonempty block, even though we have no data. This is wrong.

  • Large blocks are written (ulen (clen data)) not (ulen clen data) due to the loop in compress(). This is not the same as the format for lzop, thus a data file written using LzopCodec cannot be read by lzop. See lzop-1.03/src/p_lzo.c method lzo_compress, which contains a single very simple loop, which is how Hadoop's BlockCompressorStream should be written. This is both inefficient and wrong.

  • If the LZO compressor needs to use its holdover field (or, equivalently in other people's code, setInputFromSavedData()), then the ulen-clen-data format is broken because getBytesRead() MUST return the full number of bytes passed to setInput(), not just the number of bytes actually compressed so far; then if there is holdover data, there is nowhere for it to go but into the returned data from a second call to compress(), at which point the API has forced us to break ulen-clen-data, as per lzop's file format. This is wrong, and badly designed.

  • The number of uncompressed bytes is written to the stream in lzop. There is therefore no excuse for a "Buffer too small" error in decompression. However, this value is NOT used to resize the decompressor's output buffer, and so the error occurs. One cannot, as a rule, know the size of output buffer required to decompress a given file, so Hadoop must be configured by trial and error. This is badly designed, and harder to use.

About

Pure Java implementation of the liblzo2 LZO compression algorithm

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages