Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Go-graphsplit

A tool for splitting a large dataset into graph slices to make deals in the Filecoin Network

When storing a large dataset, we need to split it into smaller pieces to fit the sector's size, which could generally be 32GiB or 64GiB.

If we make these data into a large tarball, chunk it into small pieces, and then make storage deals with miners with these pieces, on the side of storage, it will be pretty efficient and allow us to store hundreds of TiB data in a month. However, this way will also bring difficulties for data retrieval. Even if we only needed to retrieve a small file, we would first have to retrieve and download all the pieces of this tarball, decompress it, and find the specific file we needed.

Graphsplit can solve this problem. It takes advantage of IPLD protocol, follows the Unixfs format data structures, and regards the dataset or its sub-directory as a big graph, then cuts it into small graphs. Each small graph will keep its file system structure as possible as it used to be. After that, we only need to organize these small graphs into a car file. If one data piece has a complete file and we need to retrieve it, we only need to use payload CID to retrieve it through the lotus client, fetch it back, and get the file. Besides, Graphsplit will create a manifest.csv to save the mapping with graph slice name, payload CID, Piece CID, and the inner file structure.

Another advantage of Graphsplit is it can perfectly match IPFS. Like if you build an IPFS website as your Deal UI website, the inner file structure of each data piece can be shown on it, and it is easier for users to retrieve and download the data they stored.

Build

git clone https://github.com/filedrive-team/go-graphsplit.git
cd go-graphsplit
# get submodules
git submodule update --init --recursive
# build filecoin-ffi
make ffi
make

Usage

See the work flow of graphsplit

Splitting dataset:

./graphsplit chunk \
# car-dir: folder for splitted smaller pieces, in form of .car
--car-dir=path/to/car-dir \
# slice-size: size for each pieces
--slice-size=17179869184 \
# parallel: number goroutines run when building ipld nodes
--parallel=2 \
# graph-name: it will use graph-name for prefix of smaller pieces
--graph-name=gs-test \
# calc-commp: calculation of pieceCID, default value is false. Be careful, a lot of cpu, memory and time would be consumed if slice size is very large.
--calc-commp=false \
# set true if want padding the car file to fit piece size
--add-padding=false \
# set true if want using piececid to name the chunk file
--rename=false \
# parent-path: usually just be the same as /path/to/dataset, it's just a method to figure out relative path when building IPLD graph
--parent-path=/path/to/dataset \
/path/to/dataset

Notes: A manifest.csv will created to save the mapping with graph slice name, the payload cid and slice inner structure. As following:

cat /path/to/car-dir/manifest.csv
payload_cid,filename,detail
ba...,graph-slice-name.car,inner-structure-json

If set --calc-commp=true, two another fields would be add to manifest.csv

cat /path/to/car-dir/manifest.csv
payload_cid,filename,piece_cid,piece_size,detail
ba...,graph-slice-name.car,baga...,16646144,inner-structure-json

Import car file to IPFS:

ipfs dag import /path/to/car-dir/car-file

Restore files:

# car-path: directory or file, in form of .car# output-dir: usually just be the same as /path/to/output-dir# parallel: number goroutines run when restoring
./graphsplit restore \
--car-path=/path/to/car-path \
--output-dir=/path/to/output-dir \
--parallel=2

PieceCID Calculation for a single car file:

# Calculate pieceCID for a single car file#
./graphsplit commP /path/to/carfile

Contribute

PRs are welcome!

License

MIT

About

A tool for splitting the large datasets into graph slices fit for making deals in the Filecoin Network.

Topics

Resources

Stars

45 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages