This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
This repository was archived by the owner on Aug 29, 2026. It is now read-only.

Latest commit

History

History
360 lines (241 loc) · 16.2 KB

File metadata and controls

360 lines (241 loc) · 16.2 KB
OpenKB (by PageIndex)

VectifyAI%2FOpenKB | Trendshift

OpenKB: Open LLM Knowledge Base

Scale to long documents • Reasoning-based retrieval • Native multi-modality • No Vector DB

📢 Recent Updates

  • Google Open Knowledge Format (OKF): Wiki pages follow the Google OKF specification for knowledge sharing.
  • Entity Pages: People, orgs, places, and products as dedicated wiki pages, auto-extracted and kept in sync.

📑 What is OpenKB

OpenKB (Open Knowledge Base) is an open-source system (in CLI) that compiles raw documents into a structured, interlinked wiki-style knowledge base using LLMs, powered by PageIndex's vectorless, reasoning-based retrieval for long documents.

The idea is based on a concept described by Andrej Karpathy: LLMs generate summaries, concept pages, and cross-references, all maintained automatically. Knowledge compounds over time instead of being re-derived on every query.

Why not traditional RAG?

Traditional RAG rediscovers knowledge from scratch on every query. Nothing accumulates. OpenKB compiles knowledge once into a persistent wiki, then keeps it current. Cross-references already exist, contradictions are flagged, and synthesis reflects everything consumed.

OpenKB has two layers: a wiki foundation that compiles and maintains your knowledge, and generators (query / chat / Skill Factory) that turn it into useful output. See Usage for the full command list.

Features

  • Broad format support: PDF, Word, Markdown, PowerPoint, HTML, Excel, CSV, text, URLs, and more.
  • Scales to long documents: Long and complex documents are handled via PageIndex tree indexing, enabling accurate, vectorless, context-aware retrieval.
  • Native multi-modality: Retrieves and understands figures, tables, and images, not just text.
  • Compiled wiki: The LLM compiles your documents into summaries, concept pages, entity pages, and cross-links, all kept in sync.
  • Query & chat: One-off questions or multi-turn conversations over your wiki, with persisted sessions to resume.
  • Skill Factory: Distills redistributable agent skills from your wiki.
  • OKF-ready: Wiki pages follow the Google OKF specification for knowledge sharing.
  • Obsidian-compatible: The wiki is plain .md files with cross-links. Opens in Obsidian for graph view.

🚀 Getting Started

Install

pip install openkb
Other install options:
  • Latest from GitHub:

    pip install git+https://github.com/VectifyAI/OpenKB.git
  • Install from source (editable, for development):

    git clone https://github.com/VectifyAI/OpenKB.git
    cd OpenKB
    pip install -e .

Quick Start

# 1. Create a directory for your knowledge base
mkdir my-kb &&cd my-kb
# 2. Initialize the knowledge base
openkb init
# 3. Add documents
openkb add paper.pdf
openkb add ~/papers/ # Add a whole directory
openkb add https://arxiv.org/pdf/2509.11420 # Or fetch from a URL# 4. Ask a question
openkb query "What are the main findings?"# 5. Or chat interactively
openkb chat
# (Optional) Turn the wiki into other outputs
openkb skill new my-expert "Reason like an expert on <your-topic>"# a portable agent skill
openkb visualize # an interactive knowledge graph
openkb deck new my-deck "An intro deck on <your-topic>"# slides — a single-file HTML deck

Set up your LLM

OpenKB supports multiple LLM providers (OpenAI, Claude, Gemini, etc.) via LiteLLM (pinned to a safe version).

Set your model during openkb init or in .openkb/config.yaml using the provider/model LiteLLM format (e.g. anthropic/claude-sonnet-4-6). OpenAI models can omit the prefix (e.g. gpt-5.4).

Create a .env file with your LLM API key:

LLM_API_KEY=your_llm_api_key

🧩 How OpenKB Works

Architecture

OpenKB Architecture: from raw documents (markitdown / PageIndex) through LLM wiki compilation to the wiki/ foundation, powering query/chat, the Skill Factory, and future generators

Short vs Long Document Handling

Short documentsLong documents (PDF ≥ 20 pages)
Convertmarkitdown → MarkdownPageIndex → tree index + summaries
ImagesExtracted inline (pymupdf)Extracted by PageIndex
LLM readsFull textDocument trees
Resultsummary + conceptssummary + concepts

Short documents are read in full by the LLM. Long PDFs are processed by PageIndex into a hierarchical tree index. The LLM reads the tree instead of the full text, enabling accurate and scalable retrieval for long documents.

Knowledge Compilation

When you add a document, the LLM:

  1. Generates a summary page
  2. Reads existing concept and entity pages
  3. Creates or updates concepts with cross-document synthesis
  4. Creates or updates entity pages (people, orgs, places, products)
  5. Updates the index and log

A single source might touch 10--15 wiki pages. Knowledge accumulates: each document enriches the existing wiki rather than sitting in isolation.

⚙️ Usage

OpenKB commands fall into two layers: the wiki foundation (compile + manage your knowledge) and generators (turn that wiki into useful output). Each links to a concrete walkthrough — a real artifact OpenKB generated from one sample paper (browse them all in examples/).

Layer 1: 🧱 Wiki Foundation — compile and maintain

CommandDescription
openkb initInitialize a new knowledge base (interactive)
openkb add <file_or_dir_or_URL>Add files, directories, or URLs and compile to wiki (URL content type is auto-detected)
openkb listList indexed documents and concepts
openkb statusShow knowledge base stats
openkb watchWatch raw/ and auto-compile new files
openkb lintRun structural and knowledge health checks
More wiki commands:
CommandDescription
openkb remove <doc>Remove a document and clean up its wiki pages, images, registry, and PageIndex state (--dry-run to preview, --keep-raw / --keep-empty to retain artifacts)
openkb recompile [<doc>] [--all]Re-run the compile pipeline on already-indexed docs without re-indexing. Regenerates summaries and rewrites concept pages; manual edits are overwritten (--dry-run to preview, --refresh-schema to also update wiki/AGENTS.md)
openkb feedback ["msg"]File feedback by opening a prefilled GitHub issue (--type bug/feature/question to tag it)

Example: the everyday loop walked through end to end — examples/commands/.

Layer 2: 💡 Generators — turn the wiki into output

A "generator" reads from the compiled wiki and produces something usable: an answer, a conversation, a skill folder. The wiki is the substrate; generators are the surfaces.

CommandOutputExample
openkb query "question"A grounded answer with citations (--save to persist to wiki/explorations/)query & save
openkb chatInteractive multi-turn session over the wiki (--resume, --list, --delete to manage sessions)chat
openkb visualizeA self-contained interactive knowledge graph at output/visualize/graph.html — 3D, mind-map, and radial viewsvisualize
openkb skill new <skill-name> "<intent>"Distill a redistributable agent skill from your wiki (see Skill Factory below)skills
openkb deck new <name> "<intent>"Generate a single-file HTML slide deck (--skill picks a theme, --critique runs a quality pass)slides
More skill commands:
CommandOutput
openkb skill validate [name]Validate compiled skills (auto-runs after skill new)
openkb skill eval <name>Check the skill triggers on the right prompts
openkb skill history <name> / openkb skill rollback <name>Version history + rollback for skills

🛠 Skill Factory — drop in a book; out comes a digital expert.

The flagship generator: openkb skill new distills a portable agent skill from your wiki that Claude Code, Codex, and Gemini can install and load natively. Drop in a book's worth of papers; out comes a specialist other agents can call on. → A real generated skill, plus install / share / eval / rollback, is walked through in examples/skills/.

🔧 Configuration

Settings

openkb init writes .openkb/config.yaml:

model: gpt-5.4 # LLM model (any LiteLLM-supported provider)language: en # Wiki output languagepageindex_threshold: 20# PDF pages threshold for PageIndex

The full settings reference — entity_types, OAuth providers (chatgpt/*, github_copilot/*), and LiteLLM tuning (timeouts for slow local runtimes like Ollama / LM Studio, drop_params, GitHub Copilot headers, install notes) — is in examples/configuration/.

PageIndex Setup

Long-document retrieval is a known challenge for LLMs. PageIndex solves this with vectorless, reasoning-based retrieval, by building a hierarchical tree index that lets LLMs reason over the index for context-aware retrieval.

PageIndex runs locally by default using the open-source version, with no external dependencies required.

Cloud Support(Optional):

For large or complex PDFs, PageIndex Cloud can be used to access additional capabilities, including:

  • OCR support for scanned PDFs (via hosted VLM models)
  • Faster structure generation
  • Scalable indexing for large documents

Set PAGEINDEX_API_KEY in your .env to enable cloud features:

PAGEINDEX_API_KEY=your_pageindex_api_key

Example: local vs. cloud indexing, and importing a cloud-indexed doc — examples/pageindex-cloud/.

AGENTS.md

The wiki/AGENTS.md file defines wiki structure and conventions. It's the LLM's instruction manual for maintaining the wiki. Customize it to change how your wiki is organized.

The LLM reads AGENTS.md from disk at runtime, so your edits take effect immediately.

🔌 Integrations

Using with Obsidian

The wiki is a directory of Markdown files with [[wikilinks]]. Obsidian renders it natively.

  1. Open wiki/ as an Obsidian vault
  2. Browse summaries, concepts, and explorations
  3. Use graph view to see knowledge connections
  4. Use Obsidian Web Clipper to add web articles to raw/

Using with Claude Code / Codex / Gemini CLI

OpenKB ships a SKILL.md so any agent can read your compiled wiki. No extra runtime, no MCP setup, just install the skill once.

Claude Code:
/plugin marketplace add VectifyAI/OpenKB
/plugin install openkb@vectify
OpenAI Codex CLI:

(no marketplace command yet; manual symlink)

git clone https://github.com/VectifyAI/OpenKB.git ~/openkb-src
mkdir -p ~/.agents/skills
ln -s ~/openkb-src/skills/openkb ~/.agents/skills/openkb
Gemini CLI:
gemini skills install https://github.com/VectifyAI/OpenKB.git --path skills/openkb --consent

The skill is read-only. It won't run openkb add, remove, or lint --fix without you asking. See skills/openkb/SKILL.md for the full instruction set.

🧭 Learn More

Compared to Karpathy's Approach

Karpathy's workflowOpenKB
Short documentsLLM reads directlymarkitdown → LLM reads
Long documentsContext limits, context rotPageIndex tree index
Input sourcesWeb clipper → .mdPDF, Word, PPT, Excel, HTML, text, CSV, .md, URLs
Wiki compilationLLM agentLLM agent (same)
Entity extractionManualAutomatic (people, orgs, places, products)
Q&AQuery over wikiWiki + PageIndex retrieval
OutputWiki onlyWiki + Skill Factory + agent CLI integration

The Stack

  • PageIndex — Vectorless, reasoning-based document indexing and retrieval
  • markitdown — Universal file-to-markdown conversion
  • OpenAI Agents SDK — Agent framework (supports non-OpenAI models via LiteLLM)
  • LiteLLM — Multi-provider LLM gateway
  • Click — CLI framework
  • watchdog — Filesystem monitoring

Roadmap

  • Extend long document handling to non-PDF formats
  • Scale to large document collections with nested folder support
  • Hierarchical concept (topic) indexing for massive knowledge bases
  • Database-backed storage engine
  • Web UI for browsing and managing wikis

Contributing

Contributions are welcome! Submit a pull request or open an issue for bugs and feature requests. For larger changes, consider opening an issue first to discuss the approach.

License

Apache 2.0. See LICENSE.

🌐 Open-Source Ecosystem

Other open-source projects from the PageIndex ecosystem:

  • PageIndex: Vectorless, reasoning-based RAG framework for long documents
  • ChatIndex: Tree indexing and retrieval for long conversational histories and memory
  • ConDB: A KV-cache native context database for tree-based retrieval at scale
  • PageIndex MCP: MCP server for PageIndex

Support Us

If you find OpenKB useful, please give us a star 🌟 — and check out PageIndex too!

TwitterLinkedInContact Us