Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat: add PageIndex SDK with local/cloud dual-mode support by KylinMountain · Pull Request #207 · VectifyAI/PageIndex · GitHub
Skip to content

feat: add PageIndex SDK with local/cloud dual-mode support - #207

Merged
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk
Apr 6, 2026
Merged

feat: add PageIndex SDK with local/cloud dual-mode support#207
KylinMountain merged 26 commits into
VectifyAI:devfrom
KylinMountain:feat/sdk

Conversation

@KylinMountain

Copy link
Copy Markdown
Collaborator

Summary

Unified Python SDK for document indexing and retrieval, supporting both self-hosted (local) and fully-managed (cloud) modes.

Highlights

  • Dual-mode client: LocalClient (self-hosted, user LLM key) / CloudClient (fully managed, no LLM key)
  • Collection-based multi-document management with SHA-256 dedup
  • Streaming query: col.query(stream=True) returns async-iterable QueryStream
  • Pluggable protocols: DocumentParser, StorageEngine (SQLite default)
  • Cloud backend: actual PageIndex API with SSE streaming via chat/completions

Usage

frompageindeximportLocalClient, CloudClient# Localclient=LocalClient()
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?")
# Cloudclient=CloudClient(api_key="pi-xxx")
col=client.collection()
col.add("paper.pdf")
col.query("What is this about?", stream=True)

@KylinMountain
KylinMountainforce-pushed the feat/sdk branch 4 times, most recently from f4ca4c5 to 1369cf1CompareApril 1, 2026 09:47
- Critical: preserve text in markdown structure for fallback retrieval
- Cloud: SSE response close, folder cache dict, truncate error body
- Cloud: filter internal tools, async-safe streaming via to_thread
- SQLite: multi-thread connection tracking, context manager
- Security: collection name validation, parse_pages range cap
- Polish: use count_tokens wrapper, _EXAMPLES_DIR naming, QueryStream public
- Backend protocol: add @runtime_checkable
- Replace ConfigLoader + config.yaml with Pydantic IndexConfig
- Use bool for config flags (if_add_node_summary etc.) instead of "yes"/"no"
- Enable doc_description by default for better agent QA
- Early API key validation on LocalClient init via litellm provider detection
- Expose index_config parameter on LocalClient for advanced users
- Remove config.yaml dependency from pip package
…aming
- Fix return type annotation: dict -> list (tree structure is a list)
- Fix not-found return: {} -> [] for consistency
- Cloud streaming: replace batch-then-yield with asyncio.Queue for
true real-time event delivery via background thread
…n type, legacy API fix
- Remove client-side dedup in CloudBackend (server responsibility)
- Cloud streaming: real-time via asyncio.Queue instead of batch-then-yield
- Fix get_document_structure return type: dict -> list, not-found returns []
- Fix legacy page_index() API: use IndexConfig instead of deleted ConfigLoader
- Add folder upgrade warning (once only)
- Demo: always upload, no client-side caching
KylinMountainand others added 8 commits April 3, 2026 17:27
Local demo was missing LLM provider configuration, making it fail
on first run without clear guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
….pageindex
Local-only params are now documented. Default storage_path changed from
~/.pageindex (global) to ./.pageindex (project-local) for better isolation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Was only defined on LocalClient but called from PageIndexClient._init_local(),
causing AttributeError when using PageIndexClient directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PI improvements
- Extract images from PDF pages preserving text-image reading order
using pymupdf get_text("dict") blocks. Images saved to
files/{collection}/{doc_id}/images/ with relative paths in content.
- Add get_document_structure() and get_page_content() to Collection public API
- get_document() now returns structure; add include_text param to populate
node text from page cache (WARNING in docstring: not for agent/LLM use)
- delete_document() cleans up images directory
- Agent system prompt instructs LLM to preserve image references in answers
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
KylinMountainand others added 5 commits April 5, 2026 23:59
Allows callers to specify where extracted PDF images are saved.
Default behavior unchanged (internal .pageindex/files/.../images/).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- delete_document/delete_collection now clean up custom images_dir
- add_document failure path cleans up custom images_dir
- _init_local: explicit model/retrieve_model kwargs now override
index_config dict values (was reversed)
- CloudBackend.get_document: inline structure fetch instead of calling
get_document_structure (still 2 HTTP calls but avoids method indirection)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images always stored internally at .pageindex/files/{collection}/{doc_id}/images/.
Simplifies delete/cleanup logic — no more dual-path handling.
Consumers that need images elsewhere should copy at render time.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Aligns with OpenAI Agents SDK requirement. No 3.11-specific features used.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Cloud API returns tree/ocr data in `result` field, but code only checked
`tree` and `structure` keys. Also normalizes cloud node schema
(page_index → start_index/end_index, prefix_summary → summary) and
OCR response (page_index → page, markdown → content) to match local format.
@KylinMountain
KylinMountain changed the base branch from main to devApril 6, 2026 14:47
@KylinMountain
KylinMountain marked this pull request as ready for review April 6, 2026 14:47

@claudeclaudeBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@KylinMountain