Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - SiluPanda/chunk-smart: Structure-aware text chunker for RAG pipelines · GitHub
Skip to content

Repository files navigation

chunk-smart

Structure-aware text chunker for RAG pipelines.

npm versionnpm downloadslicensenodeTypeScript

chunk-smart detects the content type of input text -- markdown, code, JSON, HTML, YAML, or plain text -- and splits it at natural structural boundaries rather than blindly by character count. Markdown is split at heading boundaries, code at function and class boundaries, JSON at top-level keys or array element groups, and plain text at paragraph and sentence boundaries. Every chunk carries rich metadata including positional offsets, token counts, heading context, and overlap indicators.

The package has zero runtime dependencies. All detection and splitting logic uses hand-written scanners and regex patterns. Token counting uses an approximate heuristic (1 token per 4 characters) suitable for most LLM tokenizers. The same input with the same options always produces the same output -- no LLM calls, no network access, fully deterministic.

Installation

npm install chunk-smart

Quick Start

import{chunk}from'chunk-smart';consttext=`# IntroductionThis is the first section with some content.## DetailsHere are the details of the implementation.## ConclusionFinal thoughts on the topic.`;constchunks=chunk(text);for(constcofchunks){console.log(c.metadata.index,c.metadata.contentType,c.metadata.tokenCount,c.content.slice(0,60));}

Output:

0 markdown 15 # Introduction\n\nThis is the first section with some c
1 markdown 14 ## Details\n\nHere are the details of the implementatio
2 markdown 10 ## Conclusion\n\nFinal thoughts on the topic.

Features

  • Auto-detection -- Identifies content type (markdown, code, JSON, HTML, YAML, plain text) from structural markers with a confidence score.
  • Structure-aware splitting -- Splits at headings, function/class boundaries, JSON keys, paragraph breaks, and sentence endings instead of arbitrary character positions.
  • Rich metadata -- Every chunk includes its sequential index, character offsets in the original text, token count, character count, content type, heading context (markdown), detected code language, and overlap indicators.
  • Configurable chunk sizing -- Control maximum and minimum token counts per chunk, with an approximate tokenizer (1 token = 4 characters).
  • Overlap support -- Configurable token-based overlap between adjacent chunks for continuity in retrieval pipelines.
  • Factory pattern -- createChunker produces a reusable chunker instance with preset defaults, avoiding repeated option parsing.
  • Type-specific methods -- Dedicated chunkMarkdown, chunkCode, and chunkJSON methods bypass auto-detection when you know the input format.
  • Zero dependencies -- No runtime dependencies. Ships as compiled CommonJS with TypeScript declarations.
  • Deterministic -- Same input and options always produce the same output. No LLM calls, no network access.

API Reference

chunk(text, options?)

Splits a text string into an array of Chunk objects. Auto-detects content type unless contentType is specified in options.

Signature:

functionchunk(text: string,options?: ChunkOptions): Chunk[]

Parameters:

ParameterTypeDescription
textstringThe text to split into chunks.
optionsChunkOptionsOptional configuration for chunking behavior.

Returns:Chunk[] -- An array of chunk objects, each containing content and metadata.

Example:

import{chunk}from'chunk-smart';// Auto-detect content typeconstchunks=chunk('# Hello\n\nWorld.\n\n## Section\n\nDetails.');// Force content type and limit chunk sizeconstsmallChunks=chunk(longText,{maxTokens: 256,contentType: 'markdown',});// Enable overlap between chunksconstoverlapping=chunk(document,{maxTokens: 512,overlap: 50,});

createChunker(defaultOptions?)

Creates a reusable chunker instance with preset default options. All methods on the returned Chunker accept optional overrides that are merged with the defaults.

Signature:

functioncreateChunker(defaultOptions?: ChunkOptions): Chunker

Parameters:

ParameterTypeDescription
defaultOptionsChunkOptionsDefault options applied to every call on the returned chunker.

Returns:Chunker -- An object with chunk, chunkMarkdown, chunkCode, chunkJSON, and detectContentType methods.

Example:

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 256,overlap: 20});constmdChunks=chunker.chunkMarkdown(markdownText);constcodeChunks=chunker.chunkCode(sourceCode);constjsonChunks=chunker.chunkJSON(jsonString);constautoChunks=chunker.chunk(unknownText);

Chunker Interface

The object returned by createChunker.

interfaceChunker{chunk(text: string,overrides?: Partial<ChunkOptions>): Chunk[];chunkMarkdown(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkCode(text: string,options?: Partial<ChunkOptions>): Chunk[];chunkJSON(text: string,options?: Partial<ChunkOptions>): Chunk[];detectContentType(text: string): DetectResult;}
MethodDescription
chunkAuto-detects content type and splits accordingly.
chunkMarkdownForces contentType: 'markdown' and splits at heading boundaries, then paragraphs, then sentences.
chunkCodeForces contentType: 'code' and splits at function, class, def, const, let, var boundaries, then blank lines.
chunkJSONForces contentType: 'json' and splits at top-level object keys or array element groups. Falls back to token-based splitting for invalid JSON.
detectContentTypeReturns the detected content type and confidence score without chunking.

detectContentType(text)

Analyzes text and returns the detected content type with a confidence score.

Signature:

functiondetectContentType(text: string): DetectResult

Parameters:

ParameterTypeDescription
textstringThe text to analyze.

Returns:DetectResult -- An object with type and confidence.

Detection heuristics:

Content TypeSignalsConfidence
jsonStarts with { or [; valid JSON.parse0.95
jsonStarts with { or [; invalid parse0.70
html<!DOCTYPE html>, <html>, or common block tags (div, p, span, body, etc.)0.90
yaml--- marker or 2+ key: value lines; no < characters0.80
markdown# headings, triple backtick fences, **bold**, list items, [link](url) (score >= 2)0.80
codefunction, class, def, import, const/let/var assignments, indented {}; lines (score >= 2)0.70
textNo structural markers detected0.50

Example:

import{detectContentType}from'chunk-smart';detectContentType('{"key": "value"}');// { type: 'json', confidence: 0.95 }detectContentType('# Title\n\nParagraph text.');// { type: 'markdown', confidence: 0.8 }detectContentType('function greet() { return "hi"; }');// { type: 'code', confidence: 0.7 }detectContentType('Just plain text with no markers.');// { type: 'text', confidence: 0.5 }

Configuration

ChunkOptions

OptionTypeDefaultDescription
maxTokensnumber512Maximum tokens per chunk. 1 token is approximately 4 characters.
minTokensnumber50Minimum tokens per chunk (informational).
overlapnumber0Number of tokens of overlap between adjacent chunks.
contentTypeContentTypeauto-detectedForce a specific content type instead of auto-detecting. One of 'markdown', 'code', 'html', 'json', 'yaml', 'text'.
preserveStructurebooleantrueWhen true, uses boundary-aware splitting (paragraphs, sentences). When false, falls back to fixed-size token splitting for html, yaml, and text types.

ContentType

typeContentType='markdown'|'code'|'html'|'json'|'yaml'|'text'

Content Type Splitting Strategies

TypePrimary BoundarySecondary BoundaryTertiary Boundary
markdown# headingsDouble newlines (paragraphs)Sentence endings
codefunction/class/def/const/let/var declarationsBlank linesCharacter-level hard split
jsonTop-level object keys or array element groupsToken-based fallback (invalid JSON)--
htmlParagraphsSentencesWord boundaries
yamlParagraphsSentencesWord boundaries
textDouble newlines (paragraphs)Sentence endings (.!?)Word boundaries

Types

Chunk

interfaceChunk{content: string;metadata: ChunkMetadata;}

ChunkMetadata

interfaceChunkMetadata{index: number;// Sequential position in the result array (0-based)startOffset: number;// Character offset of chunk start in the original textendOffset: number;// Character offset of chunk end in the original texttokenCount: number;// Approximate token count: Math.ceil(content.length / 4)charCount: number;// Character count: content.lengthcontentType: ContentType;// Detected or specified content typeheadings: string[];// Headings found within the chunk (markdown only)codeLanguage?: string;// Language detected from ``` fence or #! shebang (code only)overlapBefore: number;// Characters of overlap with the previous chunkoverlapAfter: number;// Characters of overlap with the next chunk}

DetectResult

interfaceDetectResult{type: ContentType;confidence: number;// 0.0 to 1.0}

Error Handling

chunk-smart handles edge cases gracefully without throwing exceptions:

  • Empty string -- Returns an empty array [].
  • Whitespace-only input -- Returns an empty array [].
  • Invalid JSON with contentType: 'json' -- Falls back to token-based splitting.
  • Single word exceeding maxTokens -- Hard-splits at the character limit to guarantee every chunk respects the size constraint.
  • No structural boundaries found -- Falls back to the next finer boundary level (paragraphs to sentences to words to characters).

Advanced Usage

Chunking a Codebase for Retrieval

import{createChunker}from'chunk-smart';import{readFileSync}from'node:fs';constchunker=createChunker({maxTokens: 512,overlap: 30});constsource=readFileSync('src/parser.ts','utf-8');constchunks=chunker.chunkCode(source);for(constcofchunks){console.log(`Chunk ${c.metadata.index}: ${c.metadata.tokenCount} tokens`);if(c.metadata.codeLanguage){console.log(` Language: ${c.metadata.codeLanguage}`);}}

Processing Large JSON API Responses

import{chunk}from'chunk-smart';constapiResponse=JSON.stringify(largeDataset,null,2);constchunks=chunk(apiResponse,{maxTokens: 1024,contentType: 'json',});// Each chunk is valid JSON (subset of top-level keys or array elements)for(constcofchunks){constparsed=JSON.parse(c.content);console.log(`Chunk ${c.metadata.index}: ${Object.keys(parsed).length} keys`);}

Markdown Documentation with Heading Context

import{chunk}from'chunk-smart';constdocs=readFileSync('API.md','utf-8');constchunks=chunk(docs,{maxTokens: 256,contentType: 'markdown'});for(constcofchunks){if(c.metadata.headings.length>0){console.log(`Section: ${c.metadata.headings.join(' > ')}`);}console.log(` Offset: ${c.metadata.startOffset}-${c.metadata.endOffset}`);console.log(` Tokens: ${c.metadata.tokenCount}`);}

Disabling Structure-Aware Splitting

import{chunk}from'chunk-smart';// Fixed-size token splitting without boundary awarenessconstchunks=chunk(text,{maxTokens: 128,overlap: 20,preserveStructure: false,contentType: 'text',});

Using Per-Call Overrides with a Chunker Instance

import{createChunker}from'chunk-smart';constchunker=createChunker({maxTokens: 512});// Override maxTokens for a single callconstsmall=chunker.chunk(text,{maxTokens: 128});// Detect content type without chunkingconstdetected=chunker.detectContentType(unknownInput);console.log(detected.type,detected.confidence);

TypeScript

chunk-smart is written in TypeScript with strict mode enabled. Type declarations are shipped alongside the compiled JavaScript in the dist/ directory.

All public types are exported from the package entry point:

import{chunk,createChunker,detectContentType,}from'chunk-smart';importtype{ContentType,DetectResult,ChunkMetadata,Chunk,ChunkOptions,Chunker,}from'chunk-smart';

License

MIT

About

Structure-aware text chunker for RAG pipelines

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages