Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - SiluPanda/agent-crawl: High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents. · GitHub
Skip to content

Repository files navigation

AgentCrawl

npm downloads

The High-Performance TypeScript Web Scraper for LLM Agents.

AgentCrawl is built to be the "eyes" of your AI Agents. It fetches web content, strips away the noise (ads, scripts, styles), and returns clean, token-optimized Markdown ready for your LLM context window.

It features a Hybrid Engine that starts with extremely fast static scraping and automatically falls back to a headless browser (Playwright) only when necessary for dynamic content or authentication.

Features

  • 🚀 Hybrid Engine: Instant static fetch by default, auto-switch to Headless Browser for dynamic sites.
  • Token Optimized: Returns clean Markdown, stripping 80-90% of tokens (ads, navs, footers).
  • 🧠 Agent-First: Detects Main Content, removes boilerplate, and extracts semantic structure.
  • 🔌 Plug-and-Play: Simple API designed for agent runtimes (scrape + crawl).
  • 🛡️ Production Ready: Built-in caching, retry logic, user-agent rotation, and resource blocking.
  • 🕵️ Stealth Mode: Optional best-effort browser hardening to reduce common bot-detection fingerprints.
  • Predictable Errors: Non-2xx HTTP responses are surfaced as errors instead of silently parsed as success.
  • 🇹 Type-Safe: 100% TypeScript with Zod validation.

Installation

npm install agent-crawl
# OR
bun add agent-crawl

CLI

AgentCrawl ships with a CLI for quick scraping and crawling from the terminal.

# Install globally
npm install -g agent-crawl
# Or use directly with npx
npx agent-crawl scrape https://example.com

Scrape a page to markdown

agent-crawl scrape https://example.com

JSON output with metadata

agent-crawl scrape https://example.com --output json

Browser mode for JS-rendered pages

agent-crawl scrape https://example.com --mode browser --stealth
agent-crawl scrape https://example.com --mode browser --wait-for ".content" --js "document.querySelector('.more').click()"

Structured extraction

agent-crawl scrape https://example.com --output json --extract-css '{"title":"h1","price":".price"}'
agent-crawl scrape https://example.com --output json --extract-regex '{"email":"[\\w.+-]+@[\\w-]+\\.[\\w.]+"}'

Crawl multiple pages

agent-crawl crawl https://example.com --depth 2 --pages 50 --strategy dfs
agent-crawl crawl https://example.com --robots --sitemap --include "/blog/*"
agent-crawl crawl https://example.com --strategy bestfirst --keywords "pricing,plans,features"

More options

# Auto-scroll for infinite/lazy content
agent-crawl scrape https://example.com --mode browser --scroll --max-scrolls 20
# Screenshot and PDF capture
agent-crawl scrape https://example.com --mode browser --screenshot --pdf
# Custom headers and cookies
agent-crawl scrape https://example.com -H "Authorization: Bearer tok123" --cookie "session=abc"# Table extraction and footnote-style citations
agent-crawl scrape https://example.com --tables --citations
# Pipe to an LLM
agent-crawl scrape https://docs.example.com | llm "summarize this page"

Run agent-crawl --help for the full list of options.

Quick Start

Basic Usage

import{AgentCrawl}from'agent-crawl';// Simplest usage - returns clean propertiesconstpage=awaitAgentCrawl.scrape("https://example.com");console.log(page.title);// "Example Domain"console.log(page.content);// "Example Domain\n\nThis domain is for use..."console.log(page.links);// Array of same-origin links found on the page

Advanced Usage (Optimized for LLMs)

constpage=awaitAgentCrawl.scrape("https://news.ycombinator.com",{mode: "hybrid",// "static" | "browser" | "hybrid" (default)extractMainContent: true,// Extract only the article bodyoptimizeTokens: true,// Compress excessive whitespace (default: true)stealth: true,// Enable browser stealth hardening when browser is usedstealthLevel: "balanced",// "basic" | "balanced" (default: "balanced")waitFor: ".main-content",// CSS selector to wait for (browser mode)});

Crawling Multiple Pages

Crawl an entire website with configurable depth, page limits, and concurrency:

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,// How many link-hops from start URL (default: 1)maxPages: 20,// Stop after N pages (default: 10)concurrency: 4,// Parallel requests (default: 2)extractMainContent: true,});console.log(`Crawled ${result.totalPages} pages`);console.log(`Max depth reached: ${result.maxDepthReached}`);result.pages.forEach(page=>{console.log(`- ${page.title}: ${page.url}`);});

Configuration

Scrape Options

OptionTypeDefaultDescription
mode'hybrid' | 'static' | 'browser''hybrid'Strategy to use. hybrid tries static first, then browser.
extractMainContentbooleanfalseExtract only the main article body using Readability-like algorithm.
optimizeTokensbooleantrueRemove extra whitespace and empty links for token efficiency.
stealthbooleanfalseApply best-effort browser stealth hardening (browser mode only).
stealthLevel'basic' | 'balanced''balanced'Stealth profile strength when stealth is enabled.
waitForstringundefinedCSS selector to wait for (browser mode only).
maxResponseBytesnumberundefinedBest-effort cap for static fetch response size in bytes.
httpCacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk HTTP cache for static fetch (ETag/Last-Modified).
cacheboolean | { dir?, ttlMs?, maxEntries? }undefinedOpt-in disk cache for processed scrape results (ScrapedPage).
chunkingboolean | { enabled?, maxTokens?, overlapTokens? }undefinedOpt-in token-aware chunking (page.chunks) with citation anchors.

Disk Cache Example (HTTP + Processed Result)

constpage=awaitAgentCrawl.scrape("https://example.com",{mode: "static",httpCache: {dir: ".cache/agent-crawl/http",ttlMs: 60_000,maxEntries: 1000},cache: {dir: ".cache/agent-crawl",ttlMs: 5*60_000,maxEntries: 1000},});

Chunking Example (For Agent RAG/Tools)

constpage=awaitAgentCrawl.scrape("https://example.com",{chunking: {enabled: true,maxTokens: 1200,overlapTokens: 100},});// page.chunks: [{ id, text, approxTokens, headingPath, citation: { url, anchor }}, ...]

Crawl Options

Crawl options include all scrape options plus:

OptionTypeDefaultDescription
maxDepthnumber1Maximum link depth to crawl from the start URL.
maxPagesnumber10Maximum number of pages to crawl.
concurrencynumber2Number of pages to fetch in parallel.
perHostConcurrencynumberconcurrencyMaximum concurrent requests per host.
minDelayMsnumber0Minimum delay between requests to the same host.
includePatternsstring[][]Only crawl URLs containing any of these substrings.
excludePatternsstring[][]Do not crawl URLs containing any of these substrings.
robotsboolean | { enabled?, userAgent?, respectCrawlDelay? }undefinedOpt-in robots.txt compliance (Disallow/Allow + Crawl-delay).
sitemapboolean | { enabled?, maxUrls? }undefinedOpt-in sitemap seeding from /sitemap.xml.
crawlStateboolean | { enabled?, dir?, id?, resume?, flushEvery?, persistPages? }undefinedOpt-in resumable crawl state persisted to disk.

Polite Crawl Example (Robots + Sitemap + Throttling)

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 2,maxPages: 100,concurrency: 6,perHostConcurrency: 2,minDelayMs: 250,robots: {enabled: true,userAgent: "agent-crawl",respectCrawlDelay: true},sitemap: {enabled: true,maxUrls: 1000},});

Resumable Crawl Example

constresult=awaitAgentCrawl.crawl("https://docs.example.com",{maxDepth: 3,maxPages: 500,concurrency: 6,crawlState: {enabled: true,dir: ".cache/agent-crawl/state",id: "docs-example",resume: true,flushEvery: 5,persistPages: true,},});

Return Values

scrape()ScrapedPage

{
url: string;// The final URL (after redirects)
content: string;// Clean markdown content
title?: string;// Page title
links?: string[];// Same-origin links found on the page
chunks?: Array<{// Present only when chunking is enabledid: string;text: string;approxTokens: number;headingPath: string[];citation: {url: string;anchor?: string};}>;
metadata?: {
status: number;// HTTP status code
contentLength: number;
error?: string;// Populated when scrape fails
structured?: {// Structured metadata from HTML (when present)canonicalUrl?: string;
openGraph?: Record<string,string>;
twitter?: Record<string,string>;
jsonLd?: unknown[];};// ... other headers}}

scrape() returns an empty content plus metadata.error for non-2xx responses or fetch/browser failures.

When browser rendering is used, metadata also includes:

  • stealthApplied: boolean
  • stealthLevel?: "basic" | "balanced" (when stealth is enabled)

Stealth Mode (Best-Effort)

  • Stealth is opt-in via stealth: true.
  • It is applied only for browser rendering (mode: "browser" and hybrid browser fallback).
  • It hardens common automation fingerprints (navigator.webdriver, language/plugins/platform hints, permission query behavior, and browser headers/profile).
  • It is best-effort: some anti-bot systems may still block requests.

crawl()CrawlResult

{
pages: ScrapedPage[];// Array of all scraped pages
totalPages: number;// Total number of pages crawled
maxDepthReached: number;// Deepest level reached
errors: Array<{// Any errors encounteredurl: string;error: string;}>;}

License

MIT © silupanda

About

High performance, lightweight and typesafe library to crawl and scrape web, built for LLM agents.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages