Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights

LicensePython VersionCodeQLPylintPytest

Table of Contents

Overview

Thinking Dataset leverages real-world data and efficient pipelines for business insights. We integrate robust data workflows (PDF to SQL ingestion), tactical TDD (details), and a streamlined SQLite design (details). Multi-stage AI generates STaR case studies for actionable strategies.

Features

  • 🔄 End-to-End Pipeline: Download, process, transform, and store large datasets.
  • 💾 SQLite Backend: Lightweight, fast, and easy to manage, with optional Parquet.
  • Comprehensive Testing: Thorough TDD coverage and validation checks.
  • 🖥️ Flexible CLI: Modular Click commands for quick execution of tasks.
  • 🔀 Data Transformation: Granular pipes for cleaning, merging, and deriving data.
  • 📚 STaR Case Studies: Generate synthetic scenarios alongside real data for deeper insights.
  • Parallel Execution: Efficiently process big data with optional concurrency.
  • 🔒 Secure Config: Manage environment variables discretely with .env files.

Quick Start

Prerequisites

  • Python 3.12 or later
  • Git
  • A cloud-based account (e.g., Huggingface) or a GPU (RTX 3090 or greater) for processing, or both

Setup

  1. Clone the repository:

    git clone https://github.com/MultiTonic/thinking-dataset.git
    cd thinking-dataset
  2. Install uv package manager:

    First add the package into the global environment:

    pip install uv

    Then add uv tools directory to PATH*:

    uv tool update-shell
  3. Set up the project:

    uv run setup

    *You may need to restart your terminal session for the changes to update.

This will create a virtual environment, install the project dependencies, and activate the virtual environment.

  1. Set up environment variables:

    Copy the .env.sample file to .env and change the values as needed:

    cp .env.sample .env

    Update the .env file with your credentials:

    # Required settingsHF_ORG="my_huggingface_organization"HF_USER="my_huggingface_username"HF_READ_TOKEN="my_huggingface_read_access_token"HF_WRITE_TOKEN="my_huggingface_write_access_token"# Required configurationCONFIG_PATH="config/config.yaml"# One or more providersOLLAMA_SERVER_URL="http://localhost:11434"OPENAI_API_TOKEN="your_openai_api_token"RUNPOD_API_TOKEN="your_runpod_api_token"

Usage

For complete usage instructions and examples, see the Usage Guide.

Running the Download Command

To download all parquet files from the Cablegate dataset using Hugging Face CLI:

thinking-dataset download

Running All CLI Commands

To execute all CLI commands for the project:

python assets/scripts/run_cli_commands.py

Project Structure

The following directory structure provides an overview of how the project is organized:

thinking-dataset/
├── config/ # Configuration files
├── assets/ # Assets directory for external resources
│ ├── prompts/ # Prompt templates
│ ├── scripts/ # Utility scripts
│ ├── resources/ # External project data
│ ├── templates/ # JSON prompt templates
├── data/ # Data directory
├── docs/ # Project documentation
├── reports/ # Generated reports
├── tests/ # Test files
├── thinking_dataset/ # Core project code
│ ├── commands/ # CLI command implementations
│ ├── connectors/ # Data connectors
│ ├── config/ # Configuration loaders and management
│ ├── datasets/ # Dataset definitions and processing
│ │ ├── operations/ # Data operations and transformations
│ ├── db/ # Database support
│ │ ├── operations/ # Database operations and transactions
│ ├── dto/ # Data Transfer Objects (DTO)
│ ├── io/ # File I/O operations
│ ├── pipeworks/ # Pipelines and pipes for data processing
│ │ ├── pipelines/ # Pipeline management and control
│ │ ├── pipes/ # Pipes used for data frame processing
│ ├── providers/ # AI data providers
│ ├── tonics/ # Data utility functions and helpers
│ ├── utils/ # General-purpose utility helpers
│ ├── main.py # Main execution file
└── setup.py # Project setup
└── .env # Private environment variables file

Contributing

Contributions are welcome! Fork the repository, make your changes, and create a pull request. Ensure your code follows the project's standards and includes tests. See Contributing for guidelines.

Resources

License

This dataset is licensed under the MIT License.

Citations

Please use the following BibTeX entry to cite this dataset:

@software{thinking-dataset,
author = {Kara Rawson, Joseph Pollack, and et al.},
title = {Thinking-Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation},
year = {2025},
howpublished = {\url{https://github.com/MultiTonic/thinking-dataset}},
note = {Accessed: 2025-01-25}
}

Acknowledgements

Special thanks to our contributors:

  • Kara Rawson - Lead Engineer
  • Joseph Pollack - Creator & Business Leader
  • MultiTonic Team - Support and Collaboration
  • Hugging Face - Robust tools and infrastructure for dataset management

Contact

For questions or support, please contact us at:

About

Creating a Thinking Dataset: Leveraging Real-World Data for Strategic Business Insights and STaR Case Study Generation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages