Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

DataElf

DataElf is a multi-domain analysis runtime. Its core owns job lifecycle, isolated workspaces, Pi execution, artifact validation, and finalization. Each domain owns data preparation, optional modeling, analysis instructions, output contracts, and semantic review.

CLI / API
-> JobSpec
-> typed DomainPlugin
-> domain adapter
-> optional domain modeler
-> composed Pi prompt
-> generic output validation
-> domain review

The current built-in domain is ai_index. The existing user command remains:

dataelf discover "围绕 Agentic LLMs,基于 AI Index,发现最近值得关注的 3 个 insight"

Internally this creates JobSpec(domain="ai_index", objective=...); the core does not infer a domain from arbitrary natural language.

Install

Requirements: Python 3.11+, Node.js 22.19+, and npm.

uv venv
uv pip install -e ".[dev]"
npm install

Populate the project-local Pi Fusion package declared in .pi/settings.json:

PI_CODING_AGENT_DIR=.pi/agent npm_config_cache=.pi/npm-cache \
./node_modules/.bin/pi install npm:@quarkos/pi-fusion --local --approve

Create a local configuration file:

dataelf init

dataelf.local.yaml is ignored by git and should contain local credentials.

Configuration

DataElf has one nested configuration schema. There is no flat legacy schema and no alternate explorer configuration path.

runtime:
workspace_dir: .dataelfenable_sqlite: falseexplorer:
type: pipi:
binary: ./node_modules/.bin/pimodel:
mode: jsoncwd: .timeout_seconds:
extra_args: ""log_mode: summarydomains:
ai_index:
source:
mode: apibase_url: https://index.shlab.org.cn/api/v2api_key: ak_...fixtures_dir: fixtures/ai_indexmodeling:
enabled: false# ai_index_search uses the reviewed fixed Stage 1 template.ontology_template:
stage1_config: dataelf/domains/ai_index/modeling/ontology/stage1/config.yamlstage2_config: dataelf/domains/ai_index/modeling/ontology/stage2/config.yamlraw_page_size: 50model_name:
model_max_tokens:
stage1_process_timeout_seconds: 7200stage1_request_timeout_seconds: 900stage1_request_max_retries: 3stage2_request_timeout_seconds: 600stage2_request_max_retries: 3stage2_total_timeout_seconds: 1800env:
PI_CODING_AGENT_DIR: .pi/agentOPENAI_BASE_URL: https://example.com/v1OPENAI_API_KEY: sk-...# BRAVE_API_KEY: ... # only when a Brave search skill is loaded

Configuration order:

  1. Built-in defaults.
  2. The first existing supported YAML/JSON config file.
  3. Environment variables.

Useful overrides include DATAELF_WORKSPACE, DATAELF_PI_BINARY, DATAELF_PI_MODEL, DATAELF_PI_LOG_MODE, DATAELF_AI_INDEX_MODE, DATAELF_FIXTURES_DIR, AI_INDEX_BASE_URL, AI_INDEX_API_KEY, and DATAELF_AI_INDEX_MODELING_ENABLED.

The Pi process receives only an allowlisted process environment plus the explicit env mapping and environment prepared by the selected domain adapter.

Multi-domain contract

Domain discovery is manifest-driven. A domain lives under dataelf/domains/<domain>/ and provides:

domain.yaml typed identity, plugin entrypoint, capabilities, workspace directories
plugin.py DomainPlugin implementation
config.py domain-owned typed configuration
prompt.py domain analysis method and evidence instructions
review.py domain semantic review

DomainPlugin implements these stages:

  • normalize_spec: adds deterministic domain parameters without changing the user's objective.
  • prepare: creates domain directories and prepares data access, context, environment, and input artifacts.
  • create_modeler: optionally returns a domain modeler.
  • build_prompt: supplies only domain instructions; the core composes runtime, workspace, job, artifact, and output-contract sections.
  • output_contract: declares required outputs.
  • review: applies semantic checks after generic artifact validation.
  • result_ids: exposes the domain's primary result identifiers for the workspace index.

Every stage communicates through ArtifactRef and StageResult. Core output validation checks workspace containment, required/non-empty files, JSON validity, and declared JSON roots. The final artifact_manifest.json is therefore domain-neutral.

Adding another domain does not require edits to workflow, workspace creation, prompt composition, Pi execution, or generic output validation. Tests include a fake domain with its own adapter, optional modeler, output contract, and reviewer to enforce this boundary.

AI Index domain

The AI Index plugin owns:

  • raw/ai_index/ and raw/web/ workspace directories;
  • normalized table schemas;
  • source credentials and fixture/API selection;
  • dynamic AIIndexClient access;
  • optional ontology/RDF modeling;
  • the four-phase technology-intelligence prompt;
  • insight output contracts and quality review.

Scripts written by Pi can fetch additional data through:

fromdataelf.domains.ai_index.clientimportAIIndexClientclient=AIIndexClient.from_env()
papers=client.search_papers(sub_domains=["Agentic LLMs"], page=1, size=50)

The connector supports paper, institution, and scholar search plus institution funding profiles. API responses are persisted under the job workspace and normalized into CSV tables.

Enable ontology/RDF modeling for one CLI run:

dataelf discover --ai-index-modeling \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

Use the fixed reviewed template:

dataelf discover --ai-index-modeling --ontology-template ai_index_search \
"围绕 Agentic LLMs,发现最近值得关注的 3 个 insight"

The modeler returns standard evidence artifacts; it does not replace the core prompt path. Detailed ontology operation and troubleshooting are documented in dataelf/domains/ai_index/modeling/ontology/README.md, with module responsibilities in ARCHITECTURE.md.

Workspace

Core creates only the generic skeleton; the domain adds its own directories:

.dataelf/workspaces/<job_id>/
job_spec.json
raw/
tables/
scripts/
notes/
prompts/discovery_prompt.md
logs/
reviews/quality_review.json
artifacts/
artifact_manifest.json
workspace_index.json

AI Index additionally creates raw/ai_index/, raw/web/, deep_dives/, and insights/. Final output files are not pre-created; their presence means the explorer actually produced them.

Set runtime.enable_sqlite: true only when job lookup commands are required. The workspace remains the source of job artifacts regardless of registry mode.

Pi

Project Pi settings live in .pi/settings.json, .pi/agent/models.json, and pi-harness.config.json. pi-harness.config.json belongs to @quarkos/pi-fusion and is resolved from the configured Pi working directory.

DataElf selects the Explorer model through explorer.pi.model (for example, boyuerich-openai/deepseek-v4-pro). Pi owns the provider registry, endpoint, protocol, and model metadata in .pi/agent/models.json; credentials should be supplied through the local env mapping (for example, OPENAI_API_KEY) and referenced by the provider configuration. .pi/settings.json supplies Pi's fallback provider/model only when DataElf does not pass an explicit model.

explorer.pi.log_mode can be quiet, summary, or raw. Raw JSON events are always saved to logs/pi_events.jsonl; terminal verbosity does not change the artifact contract.

Official Pi CLI resource flags belong in explorer.pi.extra_args, for example:

explorer:
type: pipi:
extra_args: "--skill /path/to/pi-skills/brave-search"

Verify

.venv/bin/python -m pytest -q
.venv/bin/python -m compileall -q dataelf

About

DataElf is an intelligent data workflow engine that turns natural-language tasks into secure, extensible, and executable data pipelines.

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages