Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

BookRAG

BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents

PaperPaper PDFCode

BookRAG is the official repository for "BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents", accepted by VLDB 2026.

BookRAG targets RAG over long, highly structured documents such as books, manuals, handbooks, reports, and technical guidebooks. Instead of flattening documents into isolated text chunks, BookRAG builds a structure-aware BookIndex that combines document hierarchies, entity relations, and fine-grained evidence mapping for more effective retrieval and generation.

News

Why BookRAG?

Complex documents are rarely just bags of chunks. They usually contain explicit logical structures, nested chapters, cross-section references, tables, figures, and entity-level dependencies. Conventional RAG pipelines often lose these signals during parsing and indexing.

BookRAG is designed around three ideas:

  • Hierarchy matters: document trees provide natural navigation paths from coarse chapters to fine evidence.
  • Relations matter: entity-level graphs capture cross-section dependencies that pure chunk retrieval can miss.
  • Query strategy matters: different questions require different retrieval workflows, from local lookup to global reasoning.

Comparison of existing RAG methods and BookRAG

Comparison of existing methods and BookRAG for complex document QA.

Method Overview

BookRAG has two stages:

  1. Offline index construction builds a BookIndex from each document.
  2. Online retrieval and generation uses the BookIndex to retrieve relevant evidence and answer user questions.

The current implementation supports multiple retrieval configurations and baselines through YAML configuration files under config and dataset configuration files under Scripts/cfg.

BookIndex Construction

BookRAG first constructs a document-native BookIndex by combining a hierarchical document tree, an entity-relation graph, and mappings between entities and document tree nodes.

BookIndex construction process

The BookIndex construction process.

Agent-based Retrieval

Given a user question, BookRAG performs query classification and planning, then dynamically composes retrieval operators over the BookIndex to locate relevant evidence and generate the final answer.

Agent-based retrieval workflow in BookRAG

The general workflow of agent-based retrieval in BookRAG.

Environment

BookRAG requires Python 3.12. The recommended setup is to create a fresh conda environment from environment.yml:

conda env create -f environment.yml
conda activate gbc-rag

Alternatively, install the Python dependencies directly:

pip install -r requirements.txt

BookRAG uses MinerU for PDF parsing and document information extraction. The default configuration uses vlm-sglang-client, so PDF parsing expects a running MinerU/SGLang service and a valid mineru.server_url in the selected YAML config. For graph extraction, the default local NLP model is en_core_web_sm, which is included in requirements.txt.

To check the installation without running models or parsing documents:

python main.py --help

If MinerU prints a warning that sglang is not installed, it can be ignored when using the client backend with an external service. Install the local SGLang runtime separately only if you plan to run MinerU's sglang-engine backend in the same environment.

Quick Start

Before running BookRAG, update the following configuration files:

Offline Index Construction

Use the example script to construct the BookIndex:

bash Scripts/example-index.sh

Online Retrieval

Run BookRAG on a configured dataset:

bash Scripts/example-rag.sh

Evaluation

We use a strong LLM as an answer extractor for evaluating model responses. Please configure the API file before evaluation, for example config/api.txt.

bash Scripts/example-eval.sh

Supported Datasets

BookRAG is evaluated on widely used complex document QA benchmarks:

This repository releases the final filtered QA files used in our experiments. The original documents, raw QA files, and dataset-specific metadata should still be obtained from the official dataset repositories above.

DatasetReleased QA fileOriginal source
MMLongBench-Docdatasets/MMLongBench-Doc.jsonMMLONGBENCH-DOC
m3docvqadatasets/m3docvqa.jsonM3DocRAG
Qasperdatasets/Qasper.jsonQasper

Each released QA file follows the unified JSON format below:

[
{
"question": "THE FIRST QUESTION",
"answer": "THE ANSWER OF FIRST QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
},
{
"question": "THE SECOND QUESTION",
"answer": "THE ANSWER OF SECOND QUESTION",
"doc_uuid": "UUID OF THE DOCUMENT PDF",
"doc_path": "PATH_TO_DIR/DOCUMENT.pdf"
}
]

Example preprocessing notebooks and scripts are available in Scripts/preprocess.

Repository Layout

BookRAG/
+-- Core/ # indexing, retrieval, generation, providers, and configs
+-- Eval/ # evaluation scripts and answer extraction utilities
+-- Scripts/ # example scripts, dataset configs, and preprocessing notebooks
+-- assets/ # figures and static assets used in the README
+-- config/ # system and method configuration files
+-- datasets/ # filtered QA files used in the paper experiments
+-- BOOKRAG_VLDB_2026_full.pdf
+-- main.py
`-- README.md

Links

License

The root LICENSE provides the GNU Affero General Public License v3.0 text for repository-level license identification. License terms for individual source files are specified by their SPDX license identifiers; the root license does not automatically apply to every file or non-software artifact in this repository.

Independently authored BookRAG source components carrying the corresponding SPDX identifier, including the Core components and the main.py program entry point, are available under either Apache-2.0 or AGPL-3.0-only. Files derived from or adapted for MinerU are available under AGPL-3.0-only. A BookRAG distribution that includes or combines these MinerU-derived components must comply with AGPL-3.0.

Datasets, evaluation resources, paper PDFs, figures, experimental artifacts, scripts and configuration files outside the licensed source components, and third-party materials are not automatically covered by the BookRAG software licenses. They remain subject to their respective terms and existing upstream notices unless explicitly stated otherwise.

The BookRAG paper has been accepted for publication in PVLDB Volume 19 and for presentation at VLDB 2026. The paper is governed separately by the PVLDB publication terms and the CC BY-NC-ND 4.0 notice stated in the paper; it is not covered by the BookRAG software licenses.

See LICENSE_SCOPE.md and THIRD_PARTY_NOTICES.md for details.

Citation

If you find BookRAG useful, please cite our paper:

@article{wang2025bookrag,
title = {BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents},
author = {Wang, Shu and Zhou, Yingli and Fang, Yixiang},
journal = {arXiv preprint arXiv:2512.03413},
year = {2025},
eprint = {2512.03413},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.03413}
}

About

No description, website, or topics provided.

Resources

Stars

129 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages