Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Wreck-It Ralph

Wreck-It Ralph

Autonomous web application security testing agent powered by Claude.

Wreck-It Ralph orchestrates Claude CLI with browser automation (Playwright MCP) to methodically test web applications for security vulnerabilities. It runs in iterations — each one a full Claude session that picks up where the last left off — with hook-based enforcement of scope, rate limits, and safety controls.

How It Works

flowchart TD
A["@targets.md + SECURITY_BRIEF.md"] --> B["Wreck-It Ralph Orchestrator"]
B --> C["Claude CLI + Playwright Browser"]
C --> D{"Testing Phase"}
D --> E["Reconnaissance"]
D --> F["Auth Testing"]
D --> G["Input Validation"]
D --> H["Access Control"]
D --> I["Business Logic"]
D --> J["API Security"]
E & F & G & H & I & J --> K["WRECK_STATUS + WRECK_FINDING + WRECK_LEARNED"]
K --> L{"More phases?"}
L -- Yes --> M["Next Iteration"]
M --> C
L -- No --> N["HTML + Markdown Reports"]
subgraph Hooks ["Safety Hooks (enforce on every action)"]
direction LR
S1["Scope Enforcer"]
S2["Rate Limiter"]
S3["Payload Validator"]
S4["Stop Validator"]
end
C -. "every tool call" .-> Hooks
Hooks -. "block or allow" .-> C
subgraph Memory ["Persisted Across Iterations"]
direction LR
M1["Learned Skills"]
M2["Findings"]
M3["Checkpoints"]
M4["Scope Learning"]
end
K --> Memory
Memory --> B
Loading

Features

Core Testing Loop

  • Phase-based testing — Reconnaissance, Authentication, Input Validation, Access Control, Business Logic, API Security
  • Iteration continuity — Context injected at each iteration start so Claude knows what was done, what's left, and what failed
  • Checkpoint recovery — Crash mid-run? Resume from the last completed iteration
  • Empty iteration detection — Exponential backoff when Claude gets stuck, auto-stops after prolonged stalling

Multi-Target Support

  • Define multiple related targets (e.g., frontend + API) in one @targets.md
  • Each target has its own scope, auth config, and type (WebApplication, Api, SinglePageApp, MobileBackend)
  • Targets can declare dependencies (DependsOn) for cross-target testing (CORS, token leakage)
  • Scope patterns are combined across all targets for the enforcer hooks

Safety Hooks (Enforced, Not Suggested)

Hooks are Node.js scripts that block Claude's actions until requirements are met. They are not prompt instructions — they are enforcement mechanisms.

HookWhat It Does
scope-enforcer.mjsBlocks navigation to out-of-scope URLs
rate-limiter.mjsEnforces requests-per-minute limit
payload-validator.mjsBlocks destructive payloads (DROP TABLE, rm -rf, etc.)
stop-validator.mjsBlocks output unless WRECK_STATUS block is present and valid
file-validator.mjsPrevents writes to wrong files
session-start.mjsInjects iteration context, skills, and blocked ops history
activity-tracker.mjsLogs all tool use for audit trail

Learned Skills System

Claude accumulates knowledge across iterations:

  • Claude-reported skills — Claude emits WRECK_LEARNED blocks when it discovers target-specific patterns (WAF behavior, auth quirks, API conventions)
  • Auto-generated failure skills — Repeated blocked operations automatically become skills so Claude stops retrying the same mistakes
  • Confidence decay — Unused skills fade over time; frequently referenced skills get boosted
  • Deduplication — Existing skills are shown to Claude with content previews to prevent redundant reports

Finding Management

  • Deduplication — Hash-based (URL + param + category + payload) and normalized title matching
  • Verification — Optional re-test of high-severity findings for confirmation
  • Evidence capture — HTTP request/response pairs and screenshots stored per finding
  • OWASP/CWE/WSTG mapping — Findings tagged with industry-standard identifiers

Scope Learning

  • Tracks repeatedly blocked hosts and suggests scope additions
  • Classifies blocked URLs by type (API endpoints, CDN, third-party services)
  • Saves suggestions to logs/scope-learning/scope-suggestions.md

Reporting

  • HTML report — Styled, self-contained report with finding details, severity breakdown, and evidence
  • Markdown report — Same content in plain text for version control or further processing
  • Generated automatically at session end (even on Ctrl+C)

Quality of Life

  • Interactive setup — Run with no arguments for a guided configuration wizard
  • System tray icon — Shows progress, current phase, finding count (Windows)
  • Audio notifications — Sounds for startup, iteration complete, finding discovered, errors
  • Toast notifications — Windows notifications for completion and errors
  • Headless mode — Run Playwright without a visible browser window
  • Temp email accounts — Auto-create test accounts via temporary email services for authenticated testing
  • Reconnaissance artifacts — Network captures, page snapshots, and screenshots preserved for review

Quick Start

# Build
dotnet build
# Run with no arguments for interactive setup
dotnet run --project src/WreckItRalph
# Or specify options directly
dotnet run --project src/WreckItRalph -- --targets @targets.md --brief SECURITY_BRIEF.md
# Validate configuration without running
dotnet run --project src/WreckItRalph -- --dry-run
# Publish self-contained binary
dotnet publish -c Release -r win-x64

CLI Options

wreck [options]
Options:
-t, --targets <file> Targets file (default: @targets.md)
-b, --brief <file> Security brief (default: SECURITY_BRIEF.md)
-m, --max-iterations <n> Max iterations (default: 50)
-d, --delay <seconds> Delay between iterations (default: 5)
--timeout <minutes> Timeout per iteration (default: 30)
--rate-limit <rpm> Requests per minute (default: 30)
--no-verify Skip finding verification
--report-dir <dir> Report output directory (default: reports)
-c, --config <file> Config file (default: wreck.json)
-s, --safe-mode Use cmd.exe without streaming output
--model <name> Claude model to use
--api-key <key> API key for the model provider
-v, --verbose Show detailed output
--no-hooks Disable safety hooks
--headless Run browser in headless mode
--dry-run Validate config only

Configuration

@targets.md

Defines testing scope, authentication, and phases.

Single target:

# Security Testing Scope## Target- Name: My Application
- Base URL: https://app.example.com- Type: WebApplication
## Authentication- Type: FormLogin
- Login URL: /login
## In-Scope-https://app.example.com/**## Out-of-Scope-https://app.example.com/admin/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Business Logic
-[ ] API Security

Multi-target:

# Security Testing Scope## Target- Name: Frontend
- Base URL: https://app.example.com- Type: SinglePageApp
- Primary: true
## Target- Name: API
- Base URL: https://api.example.com- Type: Api
- DependsOn: Frontend
## In-Scope-https://app.example.com/**-https://api.example.com/**## Testing Phases-[ ] Reconnaissance
-[ ] Authentication Testing
-[ ] Input Validation (XSS, SQLi)
-[ ] Access Control (IDOR)
-[ ] Cross-Origin Testing

SECURITY_BRIEF.md

Testing instructions and methodology for Claude. Describes the target application, known features, areas of concern, and any special testing requirements.

wreck.json (optional)

JSON configuration file for hook settings and other options:

{
"hooksConfig": {
"scopeEnforcement": true,
"rateLimiting": true,
"blockDestructive": true,
"activityTracking": true,
"contextInjection": true
}
}

Status Protocol

Claude reports status at the end of each iteration:

---WRECK_STATUS---
{"phase":"RECONNAISSANCE","status":"IN_PROGRESS","newFindings":0,"highestSeverity":"NONE","endpointsTested":5,"endpointsDiscovered":10,"exitSignal":false,"recommendation":"Continue scanning"}
---END_WRECK_STATUS---

Findings are reported inline:

---WRECK_FINDING---
{"title":"Reflected XSS in Search","severity":"HIGH","category":"XSS","url":"https://target.com/search","parameter":"q","payload":"<script>alert(1)</script>","description":"User input reflected without encoding","evidence":"Response contains unescaped payload","reproduction":"Navigate to /search, enter payload","recommendation":"HTML-encode output","cwe":"CWE-79","owasp":"A03:2021","wstg":"WSTG-INPV-01","confidence":0.9}
---END_WRECK_FINDING---

Learned skills are reported when Claude discovers reusable target-specific knowledge:

---WRECK_LEARNED---
{"skillName":"waf-blocks-inline-scripts","skillDescription":"WAF blocks script tags but allows event handlers","skillContent":"Use onerror/onload event handlers instead of <script> tags for XSS testing"}
---END_WRECK_LEARNED---

Runtime Files

When running, the tool creates:

  • wreck-hooks/ — Generated Node.js hook scripts
  • .claude/settings.local.json — Hook configuration for Claude CLI
  • logs/ — Iteration logs, context-input.json, blocked operations, learned skills
  • reports/ — Generated HTML and Markdown security reports
  • evidence/ — HTTP evidence and screenshots for findings
  • recon/ — Reconnaissance artifacts (network captures, snapshots)
  • attack-surface.md — Created by Claude during reconnaissance

Requirements

  • .NET 10.0 SDK
  • Claude CLI (claude.ai/code)
  • Node.js (for hook scripts and Playwright MCP server)

Important Notices

This tool is for authorized security testing only. You must have explicit written permission to test any target application. Unauthorized security testing is illegal in most jurisdictions.

Uses --dangerously-skip-permissions. Wreck-It Ralph runs Claude CLI with this flag to enable autonomous operation. This gives Claude unrestricted tool access within the session. The safety hooks provide guardrails, but they are not a security boundary — they are best-effort enforcement.

Scope enforcement is not airtight. Hooks validate URL patterns and payload regex, but edge cases exist. This tool assists authorized testing; it does not guarantee confinement.

Each iteration consumes Claude API credits. A typical 15-iteration run involves 15 full Claude sessions with browser automation. Monitor your usage.

Check Anthropic's acceptable use policy before using this tool for automated security testing via Claude CLI.

Project Structure

src/WreckItRalph/
├── Program.cs # CLI entry point + interactive setup
├── Config/ # WreckOptions, HooksConfig
├── Models/ # Target, Finding, WreckStatusBlock
├── Orchestration/ # Main testing loop
├── Services/ # Status parsing, findings, logging, evidence
├── Hooks/
│ ├── SafetyHookManager.cs # Hook script generation + context injection
│ ├── Scripts/ # Embedded Node.js hook scripts
│ └── Skills/ # Learned skills manager (CRUD, decay, usage)
├── Reporting/ # HTML + Markdown report generation
├── Tray/ # System tray icon + notifications
├── Setup/ # Interactive setup wizard + templates
└── Output/ # Console output formatting
tests/WreckItRalph.Tests/ # xUnit tests

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages