Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes by ChiragB254 · Pull Request #331 · VectifyAI/PageIndex · GitHub
Skip to content

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes - #331

Merged
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror
Jul 3, 2026
Merged

fix: use .get() in get_leaf_nodes to avoid KeyError on leaf nodes#331
KylinMountain merged 1 commit into
VectifyAI:devfrom
ChiragB254:fix/get-leaf-nodes-keyerror

Conversation

@ChiragB254

Copy link
Copy Markdown

Summary

  • get_leaf_nodes() crashed with KeyError: 'nodes' when called on a tree built by the standard pipeline
  • list_to_tree() calls clean_node() which deletes the nodes key entirely from leaf nodes (rather than setting it to [])
  • Accessing structure['nodes'] directly on those nodes raises KeyError
  • Fix: replace structure['nodes'] with structure.get('nodes') — returns None (falsy) safely when the key is absent

Changes

pageindex/utils.py — one-line change in get_leaf_nodes:

# Beforeifnotstructure['nodes']: # KeyError on leaf nodes# Afterifnotstructure.get('nodes'): # safe: returns None if key absent

Why this is consistent

Every other nodes access in the codebase already uses .get():

  • format_structurestructure.get('nodes')
  • is_leaf_nodenode.get('nodes')
  • print_treenode.get('nodes')

Test

frompageindex.utilsimportget_leaf_nodes, list_to_treeflat= [
{'structure': '1', 'title': 'Chapter 1', 'start_index': 1, 'end_index': 5},
{'structure': '2', 'title': 'Chapter 2', 'start_index': 6, 'end_index': 10},
]
tree=list_to_tree(flat)
leaves=get_leaf_nodes(tree) # previously raised KeyError, now returns both nodesassertlen(leaves) ==2

Fixes#330

list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ChiragB254, nice catch — consistent with the rest of the codebase. Merging! 🙏

@KylinMountain
KylinMountain merged commit cc7e43c into VectifyAI:devJul 3, 2026
GhislainAdon pushed a commit to GhislainAdon/iroko-rag that referenced this pull request Jul 6, 2026
…ctifyAI#331)
list_to_tree() deletes the 'nodes' key from leaf nodes entirely via
clean_node(). Direct access via structure['nodes'] raises KeyError on
these nodes. Using structure.get('nodes') returns None (falsy) safely,
consistent with how 'nodes' is accessed elsewhere in the codebase.
FixesVectifyAI#330
KylinMountain added a commit that referenced this pull request Jul 7, 2026
…on shims
The new SDK copied the legacy indexing pipeline into pageindex/index/
instead of moving it, leaving two divergent copies of page_index.py /
page_index_md.py / utils.py. They had already drifted (the legacy copy
still compared IndexConfig booleans against 'yes' — a separate fix),
and every pipeline change had to be applied twice.
Make pageindex/index/ the single source of truth (same pattern as the
LegacyCloudAPI shim for the 0.2.x cloud SDK):
- pageindex/index/utils.py absorbs the 27 legacy-only helpers/classes
(get_page_tokens, convert_page_to_int, ConfigLoader, PDF text helpers,
...) so it's the sole utils module. Reconciled the diverged funcs:
kept the modern versions, backported the #331 get_leaf_nodes .get()
fix, and restored remove_fields' max_len parameter (superset).
- index/page_index*.py now import `from .utils import *`;
index/legacy_utils.py (a re-export of the old top-level utils) deleted.
- Top-level page_index.py / page_index_md.py / utils.py become thin
re-export shims that emit PendingDeprecationWarning. The md_to_tree
shim coerces legacy 'yes'/'no' string flags to bool (the canonical
version is boolean-typed).
- ConfigLoader no longer reads the deleted config.yaml; it builds
defaults from IndexConfig (was an unconditional FileNotFoundError).
- __init__.py and retrieve.py import from pageindex.index.* directly so
`import pageindex` does not trip the shims.
Adds tests/test_legacy_shims.py pinning the contract: clean top-level
import doesn't warn, legacy submodule imports warn, symbols still
resolve, shim and canonical share one implementation, the #331 fix and
ConfigLoader-without-yaml both hold, and the md_to_tree coercion works.
Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ChiragB254@KylinMountain