Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - cloudigits/Krawler: Krawler - Kotlin Multiplatform Web Crawler Library · GitHub
Skip to content

Repository files navigation

Krawler 🕷️

KotlinKotlin MultiplatformLicenseMaven Central

A powerful, modern web crawling and scraping library for Kotlin Multiplatform. Build efficient web crawlers that run on JVM, Android, iOS, JavaScript, and WebAssembly with a beautiful Kotlin DSL.

✨ Features

  • 🌍 True Multiplatform: Single codebase runs on JVM, Android, iOS, JS, and WASM
  • 🎯 Intuitive Kotlin DSL: Configure crawlers with clean, type-safe syntax
  • 🚀 High Performance: Concurrent crawling with coroutines and smart rate limiting
  • 🔍 Flexible Extraction: CSS selectors, XPath, regex, and custom extractors
  • 🤖 Robots.txt Compliance: Respects website crawling policies automatically
  • 📊 Built-in Analytics: Track performance metrics and crawl statistics
  • 🔌 Extensible Architecture: Clean architecture with pluggable components
  • 💾 Smart Caching: Reduce redundant requests with intelligent caching
  • 🎨 Sample App: Full-featured Compose Multiplatform demo application

📋 Table of Contents

📦 Installation

Multiplatform Project

Add Krawler to your build.gradle.kts:

kotlin {
commonMain {
dependencies {
implementation("solutions.dreamforge.krawler:krawler:0.0.1")
}
}
}

Platform-Specific Projects

JVM/Android
dependencies {
implementation("solutions.dreamforge.krawler:krawler-jvm:0.0.1")
}
iOS
kotlin {
ios {
binaries {
framework {
baseName ="krawler"
}
}
}
}
JavaScript
dependencies {
implementation("solutions.dreamforge.krawler:krawler-js:0.0.1")
}

🚀 Quick Start

Basic Example

importsolutions.dreamforge.krawler.*importsolutions.dreamforge.krawler.dsl.*suspendfunmain() {
// Create a crawler instanceval crawler =CrawlerSDK.create()
// Define your crawl configurationval config = crawler {
name ="My First Crawler"
maxConcurrency =10
source("example") {
urls("https://example.com")
depth(2)
extract {
text("title", "h1")
text("description", "meta[name=description]")
links("links", "a[href]") {
multiple()
}
}
}
}
// Start crawling and collect results
crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> {
println("Crawled: ${result.webPage?.url}")
println("Title: ${result.webPage?.extractedData["title"]}")
}
else->println("Failed: ${result.error}")
}
}
}

Advanced Configuration

val advancedConfig = crawler {
name ="Advanced News Crawler"
maxConcurrency =20// Global extraction rules
extract {
text("title", "h1, h2, .headline") {
required()
process {
trim()
uppercase()
}
}
html("content", "article, .post-content") {
process {
// Remove ads and scripts
custom("clean-html")
}
}
// Extract structured data
regex("price", "\\$([0-9,]+\\.?[0-9]*)", group =1)
}
// Global crawl policy
policy {
respectRobotsTxt =true
delay(2000) // 2 seconds between requests
userAgent ="MyNewsBot/1.0"
maxRetries =3
timeout =15000
allowContentTypes("text/html", "application/xhtml+xml")
headers {
put("Accept-Language", "en-US,en;q=0.9")
put("Accept-Encoding", "gzip, deflate")
}
}
// Multiple sources with different configurations
source("tech-news") {
urls(
"https://techcrunch.com",
"https://theverge.com",
"https://arstechnica.com"
)
depth(3)
priority(CrawlRequest.Priority.HIGH)
// Source-specific rules
extract {
text("author", ".author-name, .by-line")
text("date", "time[datetime]")
}
}
source("business-news") {
urls("https://bloomberg.com", "https://ft.com")
depth(2)
priority(CrawlRequest.Priority.NORMAL)
policy {
delay(5000) // More conservative for premium sites
}
}
}

🔧 Platform Setup

JVM Configuration

val crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; JVM)",
maxConcurrency =50,
connectTimeoutSeconds =10,
readTimeoutSeconds =30
)
)

Android Permissions

Add to your AndroidManifest.xml:

<uses-permissionandroid:name="android.permission.INTERNET" />
<uses-permissionandroid:name="android.permission.ACCESS_NETWORK_STATE" />

iOS Configuration

No special configuration required. The library uses native iOS networking APIs.

JavaScript/Browser

// Runs in browser with CORS limitationsval crawler =CrawlerSDK.create(
SDKConfiguration(
userAgent ="MyBot/1.0 (Compatible; Browser)",
maxConcurrency =10// Limited by browser
)
)

📚 Core Concepts

Crawl Request

The fundamental unit of crawling:

val request =CrawlRequest(
id ="unique-id",
url ="https://example.com",
depth =0,
maxDepth =3,
extractionRules =listOf(/* ... */),
crawlPolicy =CrawlPolicy(/* ... */),
priority =CrawlRequest.Priority.HIGH,
metadata =mapOf("category" to "tech"),
timestamp =Clock.System.now()
)

Extraction Rules

Define what data to extract:

// CSS Selectorval titleRule =ExtractionRule(
name ="title",
selector =Selector.CssSelector("h1.main-title"),
extractionType =ExtractionType.TEXT,
required =true
)
// XPathval priceRule =ExtractionRule(
name ="price",
selector =Selector.XPathSelector("//span[@class='price']/text()"),
extractionType =ExtractionType.TEXT,
postProcessors =listOf(
PostProcessor.Extract("([0-9.]+)", 1),
PostProcessor.Custom("parse-currency")
)
)
// Regexval emailRule =ExtractionRule(
name ="emails",
selector =Selector.RegexSelector("[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}"),
extractionType =ExtractionType.TEXT,
multiple =true
)

Post Processors

Transform extracted data:

extract {
text("price", ".price") {
process {
trim()
replace("$", "")
replace(",", "")
custom("to-number")
}
}
text("description", ".desc") {
process {
trim()
substring(0, 200)
custom("remove-html") { // Configuration for custom processor
put("preserve-links", "true")
}
}
}
}

Crawl Policies

Control crawler behavior:

policy {
respectRobotsTxt =true
followRedirects =true
maxRedirects =5
delayBetweenRequests =1000// milliseconds
maxRetries =3
timeout =30000
maxContentLength =10*1024*1024// 10MB
allowContentTypes(
"text/html",
"application/xhtml+xml",
"application/xml"
)
headers {
put("Accept", "text/html,application/xhtml+xml")
put("Accept-Language", "en-US,en;q=0.9")
put("Cache-Control", "no-cache")
}
}

🔥 Advanced Usage

Batch Crawling

val requests = (1..100).map { page ->CrawlRequest(
id ="page-$page",
url ="https://example.com/products?page=$page",
// ... other configuration
)
}
crawler.batchCrawl(
requests = requests,
maxConcurrency =20,
batchId ="products-crawl"
).collect { result ->// Process results
}

Custom Extraction Engine

classMyCustomExtractor : ExtractionEngine {
overridesuspendfunextract(
html:String,
rules:List<ExtractionRule>
): Map<String, ExtractedValue> {
// Custom extraction logicreturn extractedData
}
}
val crawler =CrawlerSDK.create(
extractionEngine =MyCustomExtractor(),
// ... other components
)

Progress Monitoring

val crawler =CrawlerSDK.create()
// Monitor statistics
launch {
while (true) {
val stats = crawler.getStats()
println(""" Active: ${stats.activeCrawls} Completed: ${stats.completedCrawls} Failed: ${stats.failedCrawls} Queue Size: ${stats.queueSize} Avg Response Time: ${stats.averageResponseTime}ms""".trimIndent())
delay(1000)
}
}
// Start crawling
crawler.crawl(config).collect { /* ... */ }

Error Handling

crawler.crawl(config).collect { result ->when (result.status) {
CrawlStatus.SUCCESS-> handleSuccess(result)
CrawlStatus.ROBOTS_BLOCKED->println("Blocked by robots.txt")
CrawlStatus.TIMEOUT->println("Request timed out")
CrawlStatus.NETWORK_ERROR->println("Network error: ${result.error}")
CrawlStatus.PARSE_ERROR->println("Failed to parse: ${result.error}")
else->println("Other error: ${result.status}")
}
}

Custom Post Processors

classCurrencyParser : PostProcessorService {
overridefunregister() {
registerProcessor("parse-currency") { value, config ->val currency = config["currency"] ?:"USD"val amount = value.replace(Regex("[^0-9.]"), "").toDoubleOrNull() ?:0.0"$currency$amount"
}
}
}

🏗️ Architecture

Krawler follows Clean Architecture principles:

krawler/
├── domain/ # Business logic
│ ├── model/ # Domain models
│ ├── repository/ # Repository interfaces
│ ├── service/ # Domain services
│ └── usecase/ # Use cases
├── infrastructure/ # Implementation details
│ ├── cache/ # Caching implementation
│ ├── extraction/ # HTML parsing
│ ├── repository/ # Repository implementations
│ └── robots/ # Robots.txt handling
├── dsl/ # Kotlin DSL
├── engine/ # Crawling engine
└── http/ # HTTP client abstraction

Key Components

  • CrawlerSDK: Main entry point and facade
  • CrawlerEngine: Orchestrates crawling operations
  • ExtractionEngine: Extracts data from HTML
  • RobotsService: Handles robots.txt compliance
  • CrawlRepository: Stores crawl results
  • HttpClient: Platform-specific HTTP implementation

🎮 Sample Application

The project includes a full-featured Compose Multiplatform demo:

Running the Sample

# Desktop (JVM)
./gradlew :sample:composeApp:run
# Android# Open in Android Studio and run# iOS# Open sample/iosApp/iosApp.xcodeproj in Xcode# Web (JS)
./gradlew :sample:composeApp:jsBrowserRun
# Web (WASM)
./gradlew :sample:composeApp:wasmJsBrowserRun

Sample Features

  • Real-time crawling visualization
  • Performance metrics dashboard
  • Category-based crawling
  • Source performance tracking
  • Recent results display
  • Responsive UI for all platforms

📖 API Reference

CrawlerSDK

interfaceCrawlerSDK {
// Create crawler instancefuncreate(config:SDKConfiguration = SDKConfiguration()): CrawlerSDK// Start crawling with DSL configurationsuspendfuncrawl(configuration:CrawlerConfiguration): Flow<CrawlResult>
// Crawl single URLsuspendfuncrawlSingle(request:CrawlRequest): CrawlResult// Batch crawl multiple URLssuspendfunbatchCrawl(
requests:List<CrawlRequest>,
maxConcurrency:Int = 50,
batchId:String = "batch_${timestamp}"
): Flow<CrawlResult>
// Get statisticsfungetStats(): CrawlerStats// Stop crawlersuspendfunstop()
}

DSL Functions

// Main DSL entry pointfuncrawler(block:CrawlerConfiguration.() ->Unit): CrawlerConfiguration// Source configurationfun CrawlerConfiguration.source(name:String, block:SourceBuilder.() ->Unit)
// Extraction rulesfunextract(block:ExtractionRulesBuilder.() ->Unit)
// Crawl policyfunpolicy(block:CrawlPolicyBuilder.() ->Unit)

🤝 Contributing

We welcome contributions! Please see our Contributing Guide for details.

Development Setup

  1. Clone the repository:

    git clone https://github.com/dreamforge/krawler.git
  2. Open in IntelliJ IDEA or Android Studio

  3. Build the project:

    ./gradlew build
  4. Run tests:

    ./gradlew test

Code Style

  • Follow Kotlin coding conventions
  • Use meaningful variable and function names
  • Add KDoc comments for public APIs
  • Write unit tests for new features

📄 License

Krawler is released under the Apache License 2.0. See LICENSE for details.

Copyright 2024 DreamForge Solutions
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

🙏 Acknowledgments

📬 Contact


Made with ❤️ by DreamForge Solutions

About

Krawler - Kotlin Multiplatform Web Crawler Library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages