Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SeSEKnowing What Large Language Models Don’t Know

UAI 2026 OralPythonPyTorch

SeSE Framework Overview

Our paper 《SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory》 has been accepted by UAI 2026 as an Oral presentation (Top 1% of 1087 submissions).

Introduction

Large language models (LLMs) are increasingly deployed in safety-critical scenarios, yet they remain prone to hallucinations -- generating plausible but factually incorrect responses. Reliable uncertainty quantification (UQ) is essential for enabling LLMs to abstain from answering when uncertain, thereby mitigating these risks. We propose Semantic Structural Entropy (SeSE), a principled black-box UQ framework that works for both open- and closed-source LLMs without requiring access to internal model states.

Key Contributions

  • Principled Information Theory-based Approach: Unlike existing semantic UQ methods that overlook latent semantic structural information, SeSE reveals the intrinsic hierarchy of the LLM semantic space by constructing its optimal hierarchical abstraction based on the principle of structural entropy minimization. The structural entropy of this optimal abstraction quantifies the inherent uncertainty within the semantic space after optimal compression.

  • Theoretical Generalization: We theoretically prove that SeSE generalizes Semantic Entropy (SE), the gold standard for UQ in LLMs -- SeSE recovers SE when the encoding tree is restricted to a single layer, demonstrating it as a strictly more expressive framework.

  • Granular Claim-level Uncertainty Estimation: While existing methods focus primarily on simple short-form outputs, SeSE provides interpretable and granular claim-level uncertainty estimation for long-form generation. By constructing claim-response bipartite graphs and computing claim-level structural entropy, it captures fine-grained semantic dependencies between claims and responses.

  • State-of-the-Art Performance: Extensive experiments across 24 model-dataset combinations demonstrate SeSE is superior performance over baselines.


Project Overview

This code repository contains all the code necessary to reproduce the experiments in the paper. We have publicly released all the code and data used to generate the main experiment results.

The project consists of two main modules:

  1. Long-form uncertainty quantification -- detects hallucinations and quantifies uncertainty in paragraph-level LLM outputs via claim-response bipartite graphs
  2. Short-form generation uncertainty quantification -- quantifies semantic uncertainty in standard QA tasks (BioASQ, TriviaQA, SQuAD, etc.) via structural entropy

Project Structure

README.md Project documentation
environment.yml Conda dependencies
requirements.txt Python dependencies
long_form_structural_entropy/ Module for long-form uncertainty quantification
HCSE.py Implements Hierarchical Clustering Structural Entropy
main.py Main entry script for long-form experiments
utils.py Utility functions for evaluation
run_record/ Stores experimental outputs
sentence_structural_entropy/ Module for short-form uncertainty quantification
analyze_results.py Analyzes results and metrics
sample_answers.py Samples LLM-generated responses
uncertainty_quantification.py Implements uncertainty quantification
src/ Submodules for data processing, models, etc.
data/ Datasets
models/ Model configurations
uncertainty_measures/ UQ method implementations
utils/ Utility functions
run_record/ Stores experimental outputs

System Requirements

Hardware Dependencies

Our experiments require modern computer hardware suited for working with large language models (LLMs).

  • CPU and RAM: Intel 10th-generation CPU with 16 GB RAM or better.
  • GPU: One or more NVIDIA GPUs are required for LLM inference.
    • 7B models: NVIDIA GeForce RTX 4090 (24 GB) is sufficient.
    • 13B models: NVIDIA A100 server GPU recommended.
    • 70B models: Two NVIDIA A100 GPUs (2x80 GB) or eight RTX 4090 (8x24 GB).

Software Dependencies

  • OS: Ubuntu 20.04.6 LTS (GNU/Linux 5.15.0-89-generic x86_64)
  • Python: 3.11
  • PyTorch: 2.5.1

The file environment.yml lists the exact versions of all Python packages used in our experiments.


Installation Guide

Step 1: Install Conda

If you do not have conda installed, follow the instructions at https://conda.io/.

Step 2: Set Up Environment

conda-env update -f environment.yml
conda activate SeSE

The installation process is expected to take approximately 20 minutes.

Step 3: Set Environment Variables

  • Linux/macOS: export HUGGING_FACE_HUB_TOKEN=<your_token> export OPENAI_API_KEY=<your_api_key>

  • Windows: $env:HUGGING_FACE_HUB_TOKEN="<your_token>" $env:OPENAI_API_KEY="<your_api_key>"

Note: You may need to request access to the official Meta LLaMa model repository (apply here). The sentence-level experiments use GPT-5-Mini (OpenAI API) for accuracy assessment, which incurs variable costs (typically ~$5/run).

Datasets are automatically downloaded via Hugging Face Datasets on first execution, except for BioASQ (task b, BioASQ11, 2023), which must be manually downloaded from here and placed at ./sentence_structural_entropy/src/data/bioasq/.


Experimental Reproduction

1. Long-form Structural Entropy

Uncertainty quantification in long-form LLM-generated text.

python long_form_structural_entropy/main.py

Key Files:

  • HCSE.py -- Implements Hierarchical Clustering Structural Entropy (HCSE) calculation
  • main.py -- Main entry point (LLM calls, adjacency matrix construction, structural entropy calculation, and evaluation)
  • utils.py -- Evaluation and utility functions

Results are saved in long_form_structural_entropy/run_record/.

2. Short-form Structural Entropy

Quantifies semantic uncertainty in short-form QA tasks.

# Step 1: Sample answers
python sentence_structural_entropy/sample_answers.py --model_name=$MODEL --dataset=$DATASET$EXTRA_CFG# Step 2: Run uncertainty quantification
python sentence_structural_entropy/uncertainty_quantification.py --runid <run_id># Step 3: Analyze results
python sentence_structural_entropy/analyze_results.py --runid <run_id>

Key Files:

  • sample_answers.py -- Samples answers from LLMs
  • uncertainty_quantification.py -- Implements uncertainty quantification
  • analyze_results.py -- Analyzes uncertainty metrics (e.g., AUROC, structural entropy)
  • src/ -- Submodules for data processing, models, and utilities

For detailed parameter descriptions, refer to sentence_structural_entropy/src/utils/utils.py.

Results are saved in sentence_structural_entropy/run_record/.


Run Records

The run_record/ folders in both modules store intermediate outputs (charts, metrics, etc.) for reproducibility and analysis.


Citation

If you find this work useful in your research, please cite:

@inproceedings{
UAI2026sese,
title={Se{SE}: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory},
author={Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu},
booktitle={Forty-Second Annual Conference on Uncertainty in Artificial Intelligence},
year={2026},
url={https://openreview.net/forum?id=THZuVvy7SV}
}

About

[UAI 2026 Oral] SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory, which aims to detect hallucinated content in LLM-generated text.

Topics

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages