Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

43 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SearchFireSafety (ACL 2026)

Official dataset repository for the ACL 2026 paper: Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Overview

SearchFireSafety is a benchmark for statute-centric legal QA in the Korean fire-safety domain. The dataset is designed to evaluate:

  • Structure-aware retrieval over citation-linked legal documents
  • Multi-hop reasoning across delegated statutory provisions
  • Safe abstention behavior under partial/incomplete context

Repository Scope

This repository is organized as a dataset archive. The core release is under data/:

  • data/legal_docs.jsonl: legal corpus (article-level units) + citation links
  • data/realworld_qa.jsonl: real-world expert QA pairs
  • data/multihop_qa_mcq.jsonl: synthetic multi-hop MCQ for safety evaluation

Installation

The dataset statistics script and default graph builder use only the Python standard library. To run dense retrieval evaluation, install the minimal runtime dependencies:

pip install -r requirements.txt

If you need a CUDA-specific PyTorch build, install PyTorch for your platform first, then install the requirements above.

Dataset Statistics

Recompute these statistics with:

python scripts/compute_dataset_stats.py
CategoryStatisticNumber
Legal DocumentsTotal documents4,468
Legal DocumentsAvg. document length477.9 characters
Legal DocumentsAvg. words per document103.2
Legal DocumentsAvg. related documents1.8
Real-World Expert QATotal pairs876
Real-World Expert QAAvg. question length90.7 characters
Real-World Expert QAAvg. answer length278.1 characters
Real-World Expert QAAvg. relevant docs per question1.5
Multi-Hop QA (MCQ)Total pairs3,395
Multi-Hop QA (MCQ)Avg. question length51.1 characters
Multi-Hop QA (MCQ)Relevant docs per question2.0

File Formats

1) legal_docs.jsonl

Article-level legal corpus entries.

FieldTypeDescription
doc_idintUnique document unit ID
semantic_idstringHuman-readable legal identifier
collection_namestringParent legal collection
law_levelstringLegal hierarchy level (e.g., Act, Decree, Rule)
law_namestringLaw title
chapterstringArticle/appendix label
chapter_descriptionstringArticle heading
textstringLegal text
related_doc_idsint[]Citation/delegation-linked doc_id list

Notes:

  • 1,728 rows contain at least one outgoing related_doc_ids entry.
  • 2,740 rows have an empty related_doc_ids list.
  • related_doc_ids defines graph edges used for structure-aware retrieval.

2) realworld_qa.jsonl

Real-world public petition questions with official NFA answers.

FieldTypeDescription
question_idintQuestion ID
questionstringUser question
answerstringOfficial expert answer
related_doc_idsint[]Supporting legal document IDs
semantic_idsstring[]Supporting semantic identifiers

3) multihop_qa_mcq.jsonl

Synthetic multiple-choice QA designed to test strict multi-hop dependency.

FieldTypeDescription
question_idintQuestion ID
related_doc_idsint[]Source document IDs used to construct the question
related_semantic_idsstring[]Semantic identifiers for source docs
questionstringMCQ question
option_1 ~ option_5stringFive answer options
answer_fullint (1-5)Correct option under full context
answer_partialint (1-5)Correct option under partial context

Notes:

  • For all 3,395 rows, answer_partial = 5 ("Cannot be answered with the given information").
  • This setup explicitly evaluates safe abstention under missing evidence.

Structure-Aware Reranking Evaluation

This repository also includes lightweight scripts for evaluating Structure-Aware Reranking (SAR) on the real-world QA split.

SAR is evaluated as a strict top-100 reranker in this release. The dense retriever first retrieves the top 100 documents. SAR then keeps the same candidate set and reranks it using explicit document links from legal_docs.jsonl. The real-world QA labels are used only for evaluation, not for graph construction.

Build the Explicit Graph

Directed graph:

python scripts/build_legal_explicit_graph.py \
--output legal_explicit_graph.pkl

Undirected graph:

python scripts/build_legal_explicit_graph.py \
--undirected \
--output legal_explicit_graph_undirected.pkl

The directed graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 1,725 graph nodes with outgoing edges
  • 8,114 directed adjacency entries

The undirected graph contains:

  • 4,468 legal document nodes
  • 1,728 documents with at least one related_doc_ids entry
  • 1,650 referenced target documents
  • 2,276 graph nodes with edges
  • 15,262 undirected adjacency entries

SAR Pseudo Algorithm

Input:
q: user question
D: legal document corpus
G: explicit legal graph, where G[source_doc_id] = [neighbor_doc_id, ...]
K = 100: dense candidate size
M = 15: SAR voting seed size
F = 5: frozen dense anchor size
beta = 0.30: structural bonus weight
1. Dense retrieval
Encode q and all documents in D.
Compute dense score S_dense(d) for each document d.
Let C be the top-K documents by S_dense.
Let Seeds be the top-M documents in C.
2. Structural voting
Initialize bonus B(d) = 0 for each document d in C.
For each seed s in Seeds:
neighbors = G[s]
seed_penalty = log(|neighbors| + 1) if |neighbors| > 1 else 1
vote = S_dense(s) / seed_penalty
For each neighbor n in neighbors:
If n is not in C, skip it in strict top-100 mode.
target_penalty = log(indegree(n) + 1) if indegree(n) > 1 else 1
B(n) = B(n) + vote / target_penalty
3. Residual fusion
For each candidate d in C:
If d is one of the top-F dense anchors:
keep d above the remaining candidates.
Else:
S_SAR(d) = S_dense(d) + beta * B(d) * (1 - S_dense(d))
4. Return candidates C sorted by S_SAR.

Run Retrieval Evaluation

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100.json

For the undirected graph, replace --graph-pkl and --save-json:

python scripts/eval_realworld_graph_retrieval.py \
--graph-pkl legal_explicit_graph_undirected.pkl \
--qa-file data/realworld_qa.jsonl \
--model-name BAAI/bge-m3 \
--save-json results_legal_graph_sar_top100_undirected.json

Results

The table below reports retrieval performance on data/realworld_qa.jsonl using BAAI/bge-m3. Values are percentages. Rocchio is evaluated as a candidate-only top-100 reranker with top_k=15, alpha=0.90, and beta=0.10.

MethodR@10nDCG@10MRR@10R@20nDCG@20MRR@20R@50nDCG@50MRR@50
Baseline53.1437.3535.2661.4639.6135.8372.6642.0736.16
Rocchio52.8737.4535.3761.8739.9236.0072.7742.3136.32
SAR (Directed)55.4538.2235.5963.6240.4736.1573.4642.6536.45
SAR (Undirected)54.1837.7735.3762.9740.1835.9773.4042.4836.30

Citation

If you use this dataset, please cite the ACL 2026 paper.

@inproceedings{chae-etal-2026-evaluating,
title = "Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal {QA}",
author = "Chae, Kyubyung and Yeom, Jewon and Park, Jeongjae and Bae, Seunghyun and Jang, Ijun and Jin, Hyunbin and Jang, Jinkwan and Kim, Taesup",
editor = "Liakata, Maria and Moreira, Viviane P. and Zhang, Jiajun and Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.2112/",
doi = "10.18653/v1/2026.acl-long.2112",
pages = "45553--45573",
ISBN = "979-8-89176-390-6"
}

Contact

For questions about the dataset release, please open an issue in this repository.

About

[ACL 2026] Official dataset repository for Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages