Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Spider NL2SQL

An execution-grounded, multi-agent text-to-SQL system for the Spider benchmark. It combines hybrid schema retrieval, foreign-key graph expansion, a SQL-focused language model, structural checks, read-only SQLite execution, and automatic repair.

The original research notebooks remain in this repository. The reusable implementation now lives in an installable Python package with typed configuration, a command-line interface, and tests.

How it works

flowchart LR
Q[Question + database ID] --> R{Retrieval router}
R -->|lexical match| B[BM25 schema retrieval]
R -->|semantic match| D[E5 + FAISS retrieval]
R -->|low confidence| F[Full schema fallback]
B --> K[Foreign-key path expansion]
D --> K
K --> G[SQL generator]
F --> G
G --> S[Shape audit and optional reprompt]
S --> E[Read-only SQLite execution]
E -->|error| X[Deterministic / LLM repair]
X --> E
E --> O[SQL + JSONL trace]
Loading

The pipeline is designed around one invariant: generated SQL is never treated as successful until SQLite can execute it against the target database.

Features

  • Hybrid per-database retrieval with BM25 and E5 embeddings backed by FAISS
  • Confidence-based routing with a full-schema fallback
  • Foreign-key shortest-path expansion to include bridge tables and join edges
  • 4-bit loading for XGenerationLab/XiYanSQL-QwenCoder-7B-2502
  • English and Arabic heuristics for aggregation, grouping, uniqueness, and ranking
  • Read-only SQLite execution
  • Deterministic typo repair followed by an optional LLM repair loop
  • Batch predictions plus structured JSONL metadata and traces
  • Lightweight unit tests that do not download model weights

Repository layout

.
├── src/nl2sql/ # Installable application package
│ ├── agents.py # Generation, shape checks, execution, repair
│ ├── cli.py # `nl2sql` command
│ ├── config.py # Typed configuration
│ ├── data.py # Spider loaders
│ ├── pipeline.py # Composition and orchestration
│ ├── retrieval.py # BM25, dense retrieval, and routing
│ └── schema.py # Schema docs and foreign-key graph
├── tests/ # Fast unit tests
├── final-nl2sql-multi-agent-system.ipynb
├── Helpers/ # Research and ablation notebooks
├── Multi-agent-NL2SQL/ # Earlier multi-agent experiments
├── Dataset/ # Archived dataset artifacts
├── pyproject.toml
└── .env.example

Requirements

  • Python 3.10 or newer
  • A Spider 1.0 checkout containing tables.json and database/
  • An NVIDIA GPU is strongly recommended for the default 7B generator
  • Enough disk space for the generator and embedding model weights

The repository contains compressed dataset artifacts, not a ready-to-run Spider directory. Extract or download Spider separately and pass its root directory to the CLI.

Installation

Create an isolated environment and install the project:

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[models]"

For development tools:

python -m pip install -e ".[models,dev]"

bitsandbytes is installed only on Linux. On another platform, pass --no-4bit and ensure the model fits in available memory, or run the project in a Linux GPU environment such as Kaggle.

Dataset layout

The path supplied as --spider-root must look like this:

spider/
├── tables.json
├── dev.json
└── database/
└── concert_singer/
└── concert_singer.sqlite

An embedding model may be either a Hugging Face model ID such as intfloat/e5-small-v2 or a local Sentence Transformers checkpoint produced by the tuning notebook.

Usage

Batch prediction

nl2sql \
--spider-root /data/spider \
--embed-model intfloat/e5-small-v2 \
predict \
--input /data/spider/dev.json \
--output outputs/dev_predictions.sql \
--metadata outputs/dev_metadata.jsonl

Use --limit 10 for a smoke run. The SQL output keeps exactly one line per input example, including a safe SELECT 1 placeholder if an individual example raises an unexpected exception.

Single question

nl2sql \
--spider-root /data/spider \
query \
--db-id concert_singer \
--question "How many singers are there?"

Useful global options include --model, --embed-model, --embed-device, --no-4bit, and --no-llm-repair. Place global options before the predict or query subcommand.

Python API

frompathlibimportPathfromnl2sqlimportRuntimeConfig, build_pipelinefromnl2sql.dataimportload_spiderruntime=RuntimeConfig(
spider_root=Path("/data/spider"),
embed_model="intfloat/e5-small-v2",
)
examples, schemas, database_ids=load_spider(
runtime.spider_root/"dev.json",
runtime.tables_path,
)
pipeline=build_pipeline(runtime, schemas, database_ids)
result=pipeline.run("How many singers are there?", "concert_singer")
print(result["final_sql"])
print(result["status"], result["router_mode"])

Constructing the pipeline loads model weights and builds indexes for the supplied database IDs. For interactive workloads, construct it once and reuse it.

Outputs

The batch command creates:

  • *.sql: one predicted SQL statement per line, compatible with Spider evaluation.
  • *.jsonl: one record per example containing status, repair method, row count, router mode, remaining shape issues, runtime error, and a stage-by-stage trace.

SQLite connections use read-only mode. This prevents generated mutation statements from changing benchmark databases.

Evaluation

Use the official Spider evaluator from its repository:

git clone https://github.com/taoyds/spider.git vendor/spider
python vendor/spider/evaluation.py \
--gold /data/spider/dev_gold.sql \
--pred outputs/dev_predictions.sql \
--db /data/spider/database \
--table /data/spider/tables.json \
--etype all

Spider reports structural exact/partial matching and execution accuracy. Inspect the JSONL trace alongside aggregate metrics to distinguish retrieval, generation, shape, and execution failures.

Development

Run the fast test suite and linter before committing:

pytest
ruff check src tests

The tests exercise SQL cleanup, shape detection, deterministic repair, schema documents, and foreign-key expansion without requiring a GPU or network access.

Research notebooks

The notebooks are retained as reproducible research artifacts:

LocationPurpose
final-nl2sql-multi-agent-system.ipynbCanonical end-to-end experiment used to derive the package
Helpers/Ranker_Fine_Tuning.ipynbSpider-specific E5 retrieval tuning
Helpers/value-linker.ipynbValue-linking experiments
Helpers/hybrid-router-v2.ipynbHybrid routing experiments
Helpers/agents_without_tuning.ipynbUntuned embedding baseline
Helpers/nl2sql-rag-self-repair-hr-value-linker-agent.ipynbRAG, repair, and high-recall value-linking experiment
Multi-agent-NL2SQL/Structured-message and function-calling variants

Notebook paths still target Kaggle and may contain exploratory duplication. Use the package for repeatable runs and the notebooks for experiment history.

Additional project background is available in NL2SQL_project_documentation.pdf and Natural Language to SQL (NL2SQL).pptx.

Current limitations

  • Indexes are rebuilt in memory on each process start.
  • Shape checks are heuristic and can produce false positives.
  • Successful execution does not guarantee semantic correctness.
  • The default generator is resource-intensive and has not been abstracted behind a remote inference provider.
  • The official Spider evaluator is intentionally not vendored.

About

Multi agent system to generate SQLite queries using LLMs and RAG

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages