Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Dataset Conversion Toolkit

This toolkit provides essential utilities to process and convert various dataset formats, making them ready for machine learning training tasks. The primary features include deduplication of large datasets, conversion from Parquet to JSONL, and transformation of datasets into single-turn or multi-turn formats.

Features

  1. Deduplication (dedup.py):

    • Deduplicates large JSONL datasets in a memory-efficient manner by processing them in chunks.

    • Uses a set-based approach for efficient deduplication within each chunk.

    • drop in the folder with your jsonls and run, specify batch size according to your local requirements

      python3 dedup.py

  2. Parquet to JSONL Conversion (parquet_to_jsonl.py):

    • Converts Parquet files to JSONL format.

    • Efficiently processes large Parquet files by reading them in row groups and converting in batches.

    • drop in the folder with your parquets, specify a batch size in the script if it doesnt run

      python3 parquet_to_jsonl.py

  3. Dataset Formatting (singleormulti.py):

    • Converts datasets into either single-turn or multi-turn formats based on user input.

    • Can process datasets from various sources including local directories, individual files, or HuggingFace datasets.

    • If a dataset format is not recognized, the script prompts for input/output keys, saves the new format to the configuration, and remembers it for future use.

    • --single and --multi as args specify the output format chosen

      python3 singleormulti.py --single = {"instruction": "...", "input": "...", "output": "..."}

      python3 singleormulti.py --multi = {"conversations": [{"from": "human", "value": "..."}, {"from": "gpt", "value": "..."}, ...]}

    • --repo specifies a huggingface repo

      python3 singleormulti.py --multi --repo Open-Orca/OpenOrca

    • --dir specifies a directory (dont select the directory with the script or it will eat its own config file)

      python3 singleormulti.py --single --dir ./yourfolderwithdatasets

    • --file specifies a specific file stored locally

      python2 singleormultipy --multi --file .dataset.jsonl

Usage

Deduplication

Run dedup.py with the appropriate arguments for input, chunk size, and output directory.

Parquet to JSONL Conversion

Run parquet_to_jsonl.py with the appropriate arguments for input Parquet file and output JSONL file.

Dataset Formatting

Run singleormulti.py with the desired mode (single/multi) and specify the data source (repo, directory, or file).

Configuration

The toolkit uses a configurations.json file to understand various dataset formats. This configuration includes:

  • Identifying keys for each dataset format.
  • Classification as either single_turn or multi_turn.
  • Mappings to dictate how fields in the original dataset are converted to the desired output format.

Users can extend the configurations.json file to include new dataset formats as needed. code should detect if the dataset isnt in a recognized format and prompt the user to select the input/output tuples and will default to 3 turns of conversation in the case of turning single turn conversations to multiturn, otherwise preserves the turn number in the source dataset.

ONGOING ISSUES: does not currently detect multi turn data properly if there is more than one top level key, or the top level key has any name other than 'conversations'

singleormulti script needs cleaning, currently designed to take parquets but doesnt convert them as efficiently as parquet parquet_to_jsonl.

About

a set of scripts to easily convert all training data from huggingface into alpaca instruct or sharegpt format, which should allow for ease of use with any trainer

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages