Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

DeepSafe logo

DeepSafe-Sci

Scientific safety evaluation for language models

中文说明 · Repository · DeepSafe

Overview

DeepSafe-Sci is the standalone scientific safety evaluation repository derived from DeepSafe. It retains DeepSafe's configuration-driven inference, evaluation, metric, and reporting pipeline while adding benchmark adapters and evaluators for scientific use cases.

Scientific safety evaluation must account for two distinct failure modes. A model may provide actionable or newly synthesized information in response to a hazardous scientific request. It may also treat benign scientific work as hazardous and refuse it without adequate reason. This branch evaluates both behaviors through SciHazard, Safe-Scientist, and SOSBench.

Benchmarks

BenchmarkEvaluation focusMain signalsDefault configuration
SciHazardHarmfulness of model responses to high-risk scientific requestsExecutability level, net-new risk level, unsafe ratio, and DeHarm scoreconfigs/eval_tasks/scihazard.yaml
Safe-ScientistRecognition and safe handling of risky research tasks across scientific domainsRejection rate and safety score, reported overall and by domainconfigs/eval_tasks/safe_scientist.yaml
SOSBenchModel behavior on transformed scientific requests used to study safety and over-refusalSafe/unsafe judgment and unsafe rate, reported overall and by subjectconfigs/eval_tasks/sosbench.yaml

SciHazard data is included in this branch. It contains safe and unsafe scientific content as well as smaller subsets for development and quick checks. Safe-Scientist and SOSBench data are distributed by their respective upstream projects and must be placed locally before evaluation.

Quick start

1. Install the environment

git clone https://github.com/AI45Lab/DeepSafe-Sci.git
cd DeepSafe-Sci
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

2. Configure API access

The default benchmark configurations use an OpenAI-compatible target model and judge model. Credentials and endpoint overrides are read from environment variables rather than stored in YAML files.

export OPENAI_API_KEY="your-api-key"# Optional: set a different OpenAI-compatible endpoint.export OPENAI_BASE_URL="https://api.openai.com/v1"# Required by SciHazard when a live evidence search is needed and no cache entry is available.export SERPER_API_KEY="your-serper-api-key"

Edit the corresponding file in configs/eval_tasks/ to change the target model, judge model, concurrency, dataset path, or output directory. Both model and evaluator.judge_model_cfg should be reviewed before a run.

3. Prepare external benchmark data

SciHazard can be run with the checked-in configuration and data. For the other benchmarks:

  • Place the six Safe-Scientist domain JSON files under data/safe_scientist/.
  • Place the SOSBench Parquet files under data/sosbench/ and retain the upstream license and Responsible Use Agreement.

See docs/benchmark-data.md for the expected schemas and upstream sources.

4. Run an evaluation

Run the scripts from the repository root:

bash scripts/run_scihazard_local.sh
bash scripts/run_safe_scientist_local.sh
bash scripts/run_sosbench_local.sh

Each script passes its benchmark configuration to tools/run.py. The local runner generates model responses when predictions are absent, evaluates them with the configured judge, computes benchmark metrics, and writes a report.

5. Inspect outputs

The default output directories are:

results/scihazard/
results/safe_scientist/
results/sosbench/

A completed run normally contains:

  • predictions.jsonl: target-model responses associated with benchmark records;
  • result.json: item-level evaluation details and aggregated metrics;
  • report.md: a readable metric summary.

These outputs are runtime artifacts and are not committed to the repository.

Repository map

configs/eval_tasks/ Benchmark YAML configurations
scripts/ Repository-root launch scripts
uni_eval/datasets/ Dataset adapters
uni_eval/evaluators/ Benchmark evaluators and judges
uni_eval/metrics/ Aggregate metric implementations
DeHarmScore-trace/ SciHazard harmfulness judge and reproducibility caches
docs/benchmark-data.md External data requirements and schemas

SciHazard and DeHarmScore-trace

SciHazard uses DeHarmScore-trace as its response-harmfulness judge. In brief, the judge matches a response against a question-specific checklist, retrieves supporting evidence when needed, and assigns an Executability grade (E1-E4) and a Net-New Risk grade (N1-N4). The combined result records both the severity of operational detail and the extent to which the response contributes difficult-to-obtain information.

The bundled checklist, search-result, and search-artifact caches are specific to SciHazard and support reproducible reruns. They are not shared with Safe-Scientist or SOSBench. Prompt traces and completed model responses are intentionally excluded. See the dataset card for counts, limitations, integrity hashes, and attribution requirements. For judge-level configuration and diagnostics, see the DeHarmScore-trace quick start.

Legacy benchmark material

The repository intentionally retains older DeepSafe benchmark configurations, launch scripts, and the OpenClaw integration for reference and compatibility. Many of these files require benchmark-specific datasets, models, container images, schedulers, or environment variables and are not part of the tested SciHazard quick-start path. Internal defaults have been removed; review each legacy configuration before running it.

Responsible use, licenses, and attribution

DeepSafe-Sci is intended for controlled model evaluation and safety research. Some benchmark prompts concern hazardous scientific procedures. Run evaluations in an appropriately governed environment, restrict access to generated responses, and follow the terms of each upstream dataset.

DeepSafe-Sci software is released under Apache License 2.0. The bundled DeHarmScore-trace software component remains under its MIT license. The five SciHazard JSONL files and the bundled checklist, search-result, and search-artifact cache trees are released under CC BY 4.0; redistribution requires attribution as described in NOTICE and the dataset card. Consult each external benchmark for its additional data and usage terms.

About

No description, website, or topics provided.

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages