From a8d7d5f2f5f3c2dc8da432c5bb780e007f387a85 Mon Sep 17 00:00:00 2001 From: Pliss Boris Date: Sun, 15 Feb 2026 12:48:16 +0200 Subject: [PATCH 01/15] docs: amend constitution to v1.2.0 (add Context7 MCP plugin guidance) Add "Documentation & Library Reference" subsection under Development Workflow requiring AI assistant to use Context7 MCP plugin for up-to-date library documentation lookups. Co-Authored-By: Claude Opus 4.6 --- .specify/memory/constitution.md | 366 ++++++++++++++++++++++++++++++++ 1 file changed, 366 insertions(+) create mode 100644 .specify/memory/constitution.md diff --git a/.specify/memory/constitution.md b/.specify/memory/constitution.md new file mode 100644 index 0000000..f50bbf9 --- /dev/null +++ b/.specify/memory/constitution.md @@ -0,0 +1,366 @@ + + +# SQL Server CDC Data Extractor Constitution + +## Core Principles + +### I. Test-First / TDD (NON-NEGOTIABLE) + +Development MUST follow the Test-Driven Development cycle: + +1. **Red**: Write a failing test that defines the expected behavior. +2. **Green**: Write the minimum code to make the test pass. +3. **Refactor**: Improve code structure while keeping tests green. + +Rules: +- No production code MUST be written without a corresponding test + written first. +- Tests MUST fail before implementation begins (verify the "Red" step). +- Unit tests MUST be isolated: no database, no file system, no network. + Use interfaces/mocks/fakes for external dependencies. +- Integration tests MUST cover: IPC contracts (Named Pipes/JSON-RPC), + HTTP API client behavior, CDC interactions with SQL Server, + state store (SQL Server) read/write, configuration loading. +- Test naming convention: `MethodUnderTest_Scenario_ExpectedResult` + (e.g., `ReadCdcChanges_WhenGapDetected_ThrowsCdcGapException`). +- Test projects MUST mirror the source project structure + (e.g., `src/CdcExtractor.Core` -> `tests/CdcExtractor.Core.Tests`). +- Code coverage is a guide, not a target. Focus on meaningful + behavioral coverage, not percentage. + +### II. Domain-Driven Design (DDD) + +The codebase MUST use DDD tactical patterns where they add clarity: + +- **Bounded Contexts**: Separate the domain into clear contexts: + `Extraction` (CDC/snapshot logic), `Scheduling` (cron/triggers), + `Sink` (HTTP API client), `Configuration` (settings/state), + `Ipc` (manager-service communication), `CdcManagement` + (enable/verify/retention). +- **Entities & Value Objects**: Use Value Objects for immutable + concepts (e.g., `Lsn`, `TableIdentifier`, `BatchId`, `DatasetId`, + `SchemaHash`). Use Entities for objects with identity and lifecycle + (e.g., `TableState`, `BatchRun`). +- **Aggregates**: Group related entities under aggregate roots + (e.g., `BatchRun` owns `DatasetRun` instances). +- **Domain Events**: Use domain events for cross-context + communication (e.g., `CdcGapDetected`, `SnapshotCompleted`, + `SchemaChanged`). +- **Repository Pattern**: Abstract data access (state store, config) + behind repository interfaces. Production implementations use + SQL Server; tests use in-memory fakes. +- **Application Services**: Orchestrate use cases. Keep domain logic + in domain objects, not in services. +- **Ubiquitous Language**: Use terms from the PRD consistently: + `Batch`, `Dataset`, `Chunk`, `Snapshot`, `Delta`, `LSN`, + `Retention`, `Capture Instance`, `Re-bootstrap`. + +DDD MUST NOT be applied dogmatically. If a pattern adds complexity +without clarity (e.g., for simple CRUD or configuration), use a +simpler approach. + +### III. Exception Handling (NON-NEGOTIABLE) + +Every exception MUST be handled explicitly and logged. Silent +failures are forbidden. + +Rules: +- **No empty catch blocks**. Every `catch` MUST either: + (a) log the exception and re-throw, or + (b) log the exception and handle it with a defined recovery + strategy, or + (c) wrap the exception in a domain-specific exception with context + and throw. +- **No `catch (Exception)` without logging**. Generic catches MUST + log the full exception (message + stack trace + inner exceptions). +- **Structured exception hierarchy**: Define domain exceptions + (e.g., `CdcGapException`, `SinkUploadException`, + `PrerequisiteCheckFailedException`) that carry diagnostic context + (table name, LSN range, batch ID, etc.). +- **Fail-fast at boundaries**: Invalid configuration, missing + prerequisites, and permission errors MUST be detected at startup + or job-start and reported immediately, not mid-run. +- **Global unhandled exception handler**: The Windows Service host + and the WPF App MUST register global exception handlers + (`AppDomain.UnhandledException`, `TaskScheduler.UnobservedTaskException`, + `DispatcherUnhandledException`) that log and gracefully shut down. +- **async/await**: All `async` methods MUST propagate exceptions + correctly. Never use `async void` except for event handlers, and + those MUST have try/catch. +- **CancellationToken**: All long-running and I/O operations MUST + accept and respect `CancellationToken` for graceful shutdown. + +### IV. Observability & Logging (NON-NEGOTIABLE) + +Log messages MUST be readable and informative without requiring +knowledge of the source code. A support engineer or operator MUST +be able to understand what happened from logs alone. + +Rules: +- **Structured logging**: Use `Microsoft.Extensions.Logging` with + structured log providers (e.g., Serilog). Log entries MUST be + JSON-serializable with named properties. +- **Correlation**: Every batch run MUST have a `CorrelationId`. + Every log entry within a run MUST include `CorrelationId`, + `BatchId`, and where applicable `TableName`, `DatasetId`. +- **Log levels** MUST be used correctly: + - `Trace`: Internal diagnostic detail (CDC row-level processing). + - `Debug`: Developer-useful detail (SQL queries, HTTP requests). + - `Information`: Business events (batch started, table committed, + schema uploaded). Default production level. + - `Warning`: Recoverable issues (retry triggered, retention + raised, slow query). + - `Error`: Failures requiring attention (table skipped, HTTP 5xx, + permission denied). + - `Critical`: Service cannot continue (state corruption, global + unhandled exception). +- **Message format**: Log messages MUST follow the pattern: + `"[Action] [Subject] [Context]. [Outcome/Reason]"`. + Example: `"Uploading chunk 3/9 for table dbo.Orders + (dataset {DatasetId}). Size: {Bytes} bytes"`. + Anti-pattern: `"Error occurred"`, `"Something went wrong"`, + `"Exception in method X"`. +- **Sensitive data**: Connection strings, tokens, passwords MUST + NEVER appear in logs. Mask or redact. +- **Metrics**: Expose counters/gauges for: batch duration, rows + extracted, rows uploaded, CDC lag (LSN delta), retry count, + error count per table. +- **Windows Event Log**: Critical and Error events MUST be written + to Windows Event Log in addition to file logs, so that system + administrators can detect issues via standard monitoring. + +### V. .NET Best Practices + +The codebase MUST follow modern .NET conventions and idioms: + +- **Target framework**: .NET 8+ (LTS). Use the latest stable LTS + release available at the time of development. +- **Dependency Injection**: Use `Microsoft.Extensions.DependencyInjection` + throughout. Register services with appropriate lifetimes + (`Singleton`, `Scoped`, `Transient`). Avoid service locator + anti-pattern. +- **Configuration**: Use `Microsoft.Extensions.Configuration` + (JSON/YAML + environment variables + secrets). Bind to strongly + typed `IOptions` / `IOptionsMonitor` objects with + validation via `DataAnnotations` or `IValidateOptions`. +- **async/await**: All I/O-bound operations MUST be async. + Never use `.Result` or `.Wait()` on tasks (deadlock risk). + Use `ConfigureAwait(false)` in library code. +- **Nullable reference types**: Enable `enable` + project-wide. Treat warnings as errors for nullable. +- **Code style**: Follow `.editorconfig` rules. Use `dotnet format` + for enforcement. Naming: PascalCase for public members, `_camelCase` + for private fields, `I` prefix for interfaces. +- **Project structure**: Separate concerns into projects: + - `CdcExtractor.Domain` — domain models, interfaces, events. + - `CdcExtractor.Application` — use cases, application services. + - `CdcExtractor.Infrastructure` — SQL Server access, HTTP client, + state store, file I/O. + - `CdcExtractor.Service` — Windows Service host, scheduling, IPC. + - `CdcExtractor.App` — WPF application (configurator + manager). + - `CdcExtractor.Contracts` — shared DTOs, IPC contracts. +- **Disposable resources**: All `IDisposable`/`IAsyncDisposable` + resources MUST be disposed deterministically (`using` statements + or DI container lifetime management). +- **Immutability**: Prefer `record` types for DTOs and Value Objects. + Use `init`-only properties where mutation is not required. +- **Collections**: Expose `IReadOnlyList`, + `IReadOnlyCollection` from public APIs. Use + `ImmutableArray` for truly immutable collections where + performance matters. + +### VI. Reliability & Data Integrity + +The system MUST guarantee "at-least-once" delivery and protect +against data loss: + +- **LSN advancement**: `last_processed_lsn` MUST be updated ONLY + after successful `commit` of the corresponding dataset to + downstream. Never optimistically advance. +- **Idempotency**: All HTTP calls to downstream (chunk upload, + commit) MUST be idempotent. Use deterministic keys: + `(table, from_lsn, to_lsn, chunk_no)`. +- **CDC gap detection**: If `last_processed_lsn` is less than the + minimum available LSN, the system MUST NOT silently skip data. + It MUST log an Error, flag the table for re-bootstrap, and notify + the user via UI. +- **Transient fault handling**: Use retry with exponential backoff + for network errors, SQL deadlocks, HTTP 429/503. Classify errors + as transient vs terminal. Use Polly or equivalent. +- **Graceful shutdown**: On service stop or cancellation, the + current batch MUST be completed or cleanly aborted (not left in + an indeterminate state). +- **State store durability**: SQL Server state store MUST use + explicit transactions for LSN updates. State tables MUST reside + in the source database (or a dedicated state database on the + same instance). Use `READ COMMITTED` isolation or higher for + state read/write operations. + +### VII. Security by Default + +Security MUST be built in, not bolted on: + +- **Secrets management**: Connection strings and tokens MUST NOT be + stored in plain text config files. Use Windows Credential Manager + (DPAPI) for local secrets, environment variables for CI. +- **SQL Server access**: Use the principle of least privilege. + Document exact permissions required per operation. +- **Transport**: SQL Server connections MUST use `Encrypt=True`. + Downstream HTTP MUST use HTTPS exclusively. +- **Named Pipes ACL**: IPC pipes MUST restrict access to authorized + Windows accounts/groups only. +- **Input validation**: All external input (config values, SQL + identifiers from metadata) MUST be validated. Table/schema names + MUST be quoted with `[brackets]` to prevent SQL injection. + Parameterize all queries. +- **No secrets in logs**: Enforce via structured logging redaction + (see Principle IV). + +### VIII. Simplicity (YAGNI) + +Complexity MUST be justified. Start with the simplest solution that +meets current requirements: + +- Do NOT add features, abstractions, or configuration options that + are not required by the current phase (see `PHASES.md`). +- Prefer inline code over premature abstraction. Extract only when + duplication exceeds three occurrences or when testability demands + it. +- Prefer composition over inheritance. +- Every new project/assembly MUST have a clear, distinct + responsibility. Do NOT create projects for organizational + convenience alone. +- Configuration options MUST have sensible defaults. The user MUST + be able to get started with minimal config (connection string + + table selection). + +## Technology Stack & Constraints + +- **Runtime**: .NET 8+ (LTS), C# 12+. +- **OS**: Windows 10/11, Windows Server 2016+. +- **SQL Server**: 2016+ with CDC support (Enterprise, Standard, + or Developer edition). +- **UI**: WPF (.NET 8+) for the Windows application. +- **Service host**: `Microsoft.Extensions.Hosting` with + `BackgroundService` registered as a Windows Service via + `Microsoft.Extensions.Hosting.WindowsServices`. +- **IPC**: Windows Named Pipes with StreamJsonRpc (Microsoft). +- **HTTP client**: `HttpClient` via `IHttpClientFactory` with + Polly for resilience policies. +- **State store**: SQL Server (same instance as source data or + dedicated). Access via `Microsoft.Data.SqlClient` + Dapper or + Entity Framework Core (choose one per project convention). +- **Logging**: `Microsoft.Extensions.Logging` + Serilog + (file sink + Windows Event Log sink). +- **Testing**: xUnit + Moq (or NSubstitute) + FluentAssertions. + Integration tests use Testcontainers for SQL Server where + feasible. +- **Build**: `dotnet build` / `dotnet test` / `dotnet publish`. + MSBuild for packaging. +- **Code quality**: `.editorconfig`, `dotnet format`, + nullable reference types enabled, warnings-as-errors in CI. +- **Secrets**: DPAPI (Windows Credential Manager) for local + token/credential storage. + +## Development Workflow & Quality Gates + +### Workflow + +1. **Branch per feature/fix**: Create a branch from `main`. + Naming: `/` + (e.g., `feature/cdc-gap-detection`, `fix/retry-deadlock`). +2. **TDD cycle**: For every change: + - Write/update tests first (Red). + - Implement (Green). + - Refactor. + - Commit with passing tests. +3. **Small, focused commits**: Each commit MUST compile and pass + all tests. Commit message format: + `(): ` (Conventional Commits). +4. **Pull request**: All changes MUST go through PR with review. + PR description MUST reference the spec/task. + +### Documentation & Library Reference + +When implementing features that use external libraries or frameworks, +the AI coding assistant MUST use the **Context7 MCP plugin** to +retrieve up-to-date documentation and code examples. This ensures +implementations follow current API conventions rather than relying +on potentially outdated training data. + +Rules: +- **Before using an unfamiliar API**: The assistant MUST query + Context7 (`resolve-library-id` then `query-docs`) to obtain + current documentation for the target library. +- **When debugging library-related issues**: The assistant SHOULD + query Context7 to verify correct API usage against the latest + docs before proposing fixes. +- **Primary project dependencies** covered by this guidance: + StreamJsonRpc, Dapper, Serilog, Polly, xUnit, + FluentAssertions, Moq/NSubstitute, Testcontainers, + Microsoft.Extensions.* (DI, Configuration, Hosting, Logging). +- **When Context7 has no results**: Fall back to official + documentation websites. Document the lookup attempt so future + sessions know the library is not indexed. +- **Do NOT query Context7** for .NET BCL (Base Class Library) or + C# language features — these are well-known and stable. + +### Quality Gates (all MUST pass before merge) + +1. **Build**: `dotnet build` succeeds with zero warnings + (warnings-as-errors). +2. **Tests**: All unit and integration tests pass. +3. **Code format**: `dotnet format --verify-no-changes` passes. +4. **No TODOs without tracking**: Any `TODO` in code MUST + reference a task/issue ID. +5. **Constitution compliance**: Reviewer MUST verify: + - No silent exceptions (Principle III). + - Log messages are informative (Principle IV). + - New code follows DDD where applicable (Principle II). + - No unnecessary complexity (Principle VIII). + +## Governance + +This constitution is the authoritative source for development +standards in the SQL Server CDC Data Extractor project. It +supersedes informal practices and ad-hoc conventions. + +### Amendment Procedure + +1. Propose a change via PR modifying this file. +2. Describe the rationale and impact in the PR description. +3. Update `CONSTITUTION_VERSION` following semantic versioning: + - **MAJOR**: Principle removed, redefined, or made incompatible. + - **MINOR**: New principle/section added or materially expanded. + - **PATCH**: Clarifications, wording fixes, non-semantic changes. +4. Update `LAST_AMENDED_DATE` to the merge date. +5. Propagate changes to dependent templates if affected. + +### Compliance + +- All PRs and code reviews MUST verify compliance with this + constitution. +- Deviations MUST be documented and justified in the PR with a + reference to the specific principle being deviated from and the + reason. +- Runtime development guidance (prompts, agent instructions) MUST + reference this constitution for principle enforcement. + +**Version**: 1.2.0 | **Ratified**: 2026-02-15 | **Last Amended**: 2026-02-15 From f79dc202ddf46545da27038dda147853d9c76a4d Mon Sep 17 00:00:00 2001 From: Pliss Boris Date: Sun, 15 Feb 2026 12:48:48 +0200 Subject: [PATCH 02/15] chore: add speckit config, templates, and MVP feature specs Add Claude commands, speckit templates, PowerShell setup scripts, CLAUDE.md project guidelines, and 001-mvp-end-to-end feature specification with plan, research, data model, contracts, and tasks. Co-Authored-By: Claude Opus 4.6 --- .claude/commands/speckit.analyze.md | 184 ++++++ .claude/commands/speckit.checklist.md | 294 +++++++++ .claude/commands/speckit.clarify.md | 181 ++++++ .claude/commands/speckit.constitution.md | 84 +++ .claude/commands/speckit.implement.md | 135 ++++ .claude/commands/speckit.plan.md | 89 +++ .claude/commands/speckit.specify.md | 258 ++++++++ .claude/commands/speckit.tasks.md | 137 ++++ .claude/commands/speckit.taskstoissues.md | 30 + .claude/settings.local.json | 11 + .../powershell/check-prerequisites.ps1 | 148 +++++ .specify/scripts/powershell/common.ps1 | 137 ++++ .../scripts/powershell/create-new-feature.ps1 | 283 ++++++++ .specify/scripts/powershell/setup-plan.ps1 | 61 ++ .../powershell/update-agent-context.ps1 | 451 +++++++++++++ .specify/templates/agent-file-template.md | 28 + .specify/templates/checklist-template.md | 40 ++ .specify/templates/constitution-template.md | 50 ++ .specify/templates/plan-template.md | 104 +++ .specify/templates/spec-template.md | 115 ++++ .specify/templates/tasks-template.md | 251 +++++++ CLAUDE.md | 29 + .../checklists/requirements.md | 41 ++ .../contracts/downstream-api-client.md | 279 ++++++++ .../contracts/ipc-contract.md | 279 ++++++++ specs/001-mvp-end-to-end/data-model.md | 301 +++++++++ specs/001-mvp-end-to-end/plan.md | 288 ++++++++ specs/001-mvp-end-to-end/quickstart.md | 156 +++++ specs/001-mvp-end-to-end/research.md | 165 +++++ specs/001-mvp-end-to-end/spec.md | 615 ++++++++++++++++++ specs/001-mvp-end-to-end/tasks.md | 425 ++++++++++++ 31 files changed, 5649 insertions(+) create mode 100644 .claude/commands/speckit.analyze.md create mode 100644 .claude/commands/speckit.checklist.md create mode 100644 .claude/commands/speckit.clarify.md create mode 100644 .claude/commands/speckit.constitution.md create mode 100644 .claude/commands/speckit.implement.md create mode 100644 .claude/commands/speckit.plan.md create mode 100644 .claude/commands/speckit.specify.md create mode 100644 .claude/commands/speckit.tasks.md create mode 100644 .claude/commands/speckit.taskstoissues.md create mode 100644 .claude/settings.local.json create mode 100644 .specify/scripts/powershell/check-prerequisites.ps1 create mode 100644 .specify/scripts/powershell/common.ps1 create mode 100644 .specify/scripts/powershell/create-new-feature.ps1 create mode 100644 .specify/scripts/powershell/setup-plan.ps1 create mode 100644 .specify/scripts/powershell/update-agent-context.ps1 create mode 100644 .specify/templates/agent-file-template.md create mode 100644 .specify/templates/checklist-template.md create mode 100644 .specify/templates/constitution-template.md create mode 100644 .specify/templates/plan-template.md create mode 100644 .specify/templates/spec-template.md create mode 100644 .specify/templates/tasks-template.md create mode 100644 CLAUDE.md create mode 100644 specs/001-mvp-end-to-end/checklists/requirements.md create mode 100644 specs/001-mvp-end-to-end/contracts/downstream-api-client.md create mode 100644 specs/001-mvp-end-to-end/contracts/ipc-contract.md create mode 100644 specs/001-mvp-end-to-end/data-model.md create mode 100644 specs/001-mvp-end-to-end/plan.md create mode 100644 specs/001-mvp-end-to-end/quickstart.md create mode 100644 specs/001-mvp-end-to-end/research.md create mode 100644 specs/001-mvp-end-to-end/spec.md create mode 100644 specs/001-mvp-end-to-end/tasks.md diff --git a/.claude/commands/speckit.analyze.md b/.claude/commands/speckit.analyze.md new file mode 100644 index 0000000..542a3de --- /dev/null +++ b/.claude/commands/speckit.analyze.md @@ -0,0 +1,184 @@ +--- +description: Perform a non-destructive cross-artifact consistency and quality analysis across spec.md, plan.md, and tasks.md after task generation. +--- + +## User Input + +```text +$ARGUMENTS +``` + +You **MUST** consider the user input before proceeding (if not empty). + +## Goal + +Identify inconsistencies, duplications, ambiguities, and underspecified items across the three core artifacts (`spec.md`, `plan.md`, `tasks.md`) before implementation. This command MUST run only after `/speckit.tasks` has successfully produced a complete `tasks.md`. + +## Operating Constraints + +**STRICTLY READ-ONLY**: Do **not** modify any files. Output a structured analysis report. Offer an optional remediation plan (user must explicitly approve before any follow-up editing commands would be invoked manually). + +**Constitution Authority**: The project constitution (`.specify/memory/constitution.md`) is **non-negotiable** within this analysis scope. Constitution conflicts are automatically CRITICAL and require adjustment of the spec, plan, or tasks—not dilution, reinterpretation, or silent ignoring of the principle. If a principle itself needs to change, that must occur in a separate, explicit constitution update outside `/speckit.analyze`. + +## Execution Steps + +### 1. Initialize Analysis Context + +Run `.specify/scripts/powershell/check-prerequisites.ps1 -Json -RequireTasks -IncludeTasks` once from repo root and parse JSON for FEATURE_DIR and AVAILABLE_DOCS. Derive absolute paths: + +- SPEC = FEATURE_DIR/spec.md +- PLAN = FEATURE_DIR/plan.md +- TASKS = FEATURE_DIR/tasks.md + +Abort with an error message if any required file is missing (instruct the user to run missing prerequisite command). +For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot"). + +### 2. Load Artifacts (Progressive Disclosure) + +Load only the minimal necessary context from each artifact: + +**From spec.md:** + +- Overview/Context +- Functional Requirements +- Non-Functional Requirements +- User Stories +- Edge Cases (if present) + +**From plan.md:** + +- Architecture/stack choices +- Data Model references +- Phases +- Technical constraints + +**From tasks.md:** + +- Task IDs +- Descriptions +- Phase grouping +- Parallel markers [P] +- Referenced file paths + +**From constitution:** + +- Load `.specify/memory/constitution.md` for principle validation + +### 3. Build Semantic Models + +Create internal representations (do not include raw artifacts in output): + +- **Requirements inventory**: Each functional + non-functional requirement with a stable key (derive slug based on imperative phrase; e.g., "User can upload file" → `user-can-upload-file`) +- **User story/action inventory**: Discrete user actions with acceptance criteria +- **Task coverage mapping**: Map each task to one or more requirements or stories (inference by keyword / explicit reference patterns like IDs or key phrases) +- **Constitution rule set**: Extract principle names and MUST/SHOULD normative statements + +### 4. Detection Passes (Token-Efficient Analysis) + +Focus on high-signal findings. Limit to 50 findings total; aggregate remainder in overflow summary. + +#### A. Duplication Detection + +- Identify near-duplicate requirements +- Mark lower-quality phrasing for consolidation + +#### B. Ambiguity Detection + +- Flag vague adjectives (fast, scalable, secure, intuitive, robust) lacking measurable criteria +- Flag unresolved placeholders (TODO, TKTK, ???, ``, etc.) + +#### C. Underspecification + +- Requirements with verbs but missing object or measurable outcome +- User stories missing acceptance criteria alignment +- Tasks referencing files or components not defined in spec/plan + +#### D. Constitution Alignment + +- Any requirement or plan element conflicting with a MUST principle +- Missing mandated sections or quality gates from constitution + +#### E. Coverage Gaps + +- Requirements with zero associated tasks +- Tasks with no mapped requirement/story +- Non-functional requirements not reflected in tasks (e.g., performance, security) + +#### F. Inconsistency + +- Terminology drift (same concept named differently across files) +- Data entities referenced in plan but absent in spec (or vice versa) +- Task ordering contradictions (e.g., integration tasks before foundational setup tasks without dependency note) +- Conflicting requirements (e.g., one requires Next.js while other specifies Vue) + +### 5. Severity Assignment + +Use this heuristic to prioritize findings: + +- **CRITICAL**: Violates constitution MUST, missing core spec artifact, or requirement with zero coverage that blocks baseline functionality +- **HIGH**: Duplicate or conflicting requirement, ambiguous security/performance attribute, untestable acceptance criterion +- **MEDIUM**: Terminology drift, missing non-functional task coverage, underspecified edge case +- **LOW**: Style/wording improvements, minor redundancy not affecting execution order + +### 6. Produce Compact Analysis Report + +Output a Markdown report (no file writes) with the following structure: + +## Specification Analysis Report + +| ID | Category | Severity | Location(s) | Summary | Recommendation | +|----|----------|----------|-------------|---------|----------------| +| A1 | Duplication | HIGH | spec.md:L120-134 | Two similar requirements ... | Merge phrasing; keep clearer version | + +(Add one row per finding; generate stable IDs prefixed by category initial.) + +**Coverage Summary Table:** + +| Requirement Key | Has Task? | Task IDs | Notes | +|-----------------|-----------|----------|-------| + +**Constitution Alignment Issues:** (if any) + +**Unmapped Tasks:** (if any) + +**Metrics:** + +- Total Requirements +- Total Tasks +- Coverage % (requirements with >=1 task) +- Ambiguity Count +- Duplication Count +- Critical Issues Count + +### 7. Provide Next Actions + +At end of report, output a concise Next Actions block: + +- If CRITICAL issues exist: Recommend resolving before `/speckit.implement` +- If only LOW/MEDIUM: User may proceed, but provide improvement suggestions +- Provide explicit command suggestions: e.g., "Run /speckit.specify with refinement", "Run /speckit.plan to adjust architecture", "Manually edit tasks.md to add coverage for 'performance-metrics'" + +### 8. Offer Remediation + +Ask the user: "Would you like me to suggest concrete remediation edits for the top N issues?" (Do NOT apply them automatically.) + +## Operating Principles + +### Context Efficiency + +- **Minimal high-signal tokens**: Focus on actionable findings, not exhaustive documentation +- **Progressive disclosure**: Load artifacts incrementally; don't dump all content into analysis +- **Token-efficient output**: Limit findings table to 50 rows; summarize overflow +- **Deterministic results**: Rerunning without changes should produce consistent IDs and counts + +### Analysis Guidelines + +- **NEVER modify files** (this is read-only analysis) +- **NEVER hallucinate missing sections** (if absent, report them accurately) +- **Prioritize constitution violations** (these are always CRITICAL) +- **Use examples over exhaustive rules** (cite specific instances, not generic patterns) +- **Report zero issues gracefully** (emit success report with coverage statistics) + +## Context + +$ARGUMENTS diff --git a/.claude/commands/speckit.checklist.md b/.claude/commands/speckit.checklist.md new file mode 100644 index 0000000..b15f916 --- /dev/null +++ b/.claude/commands/speckit.checklist.md @@ -0,0 +1,294 @@ +--- +description: Generate a custom checklist for the current feature based on user requirements. +--- + +## Checklist Purpose: "Unit Tests for English" + +**CRITICAL CONCEPT**: Checklists are **UNIT TESTS FOR REQUIREMENTS WRITING** - they validate the quality, clarity, and completeness of requirements in a given domain. + +**NOT for verification/testing**: + +- ❌ NOT "Verify the button clicks correctly" +- ❌ NOT "Test error handling works" +- ❌ NOT "Confirm the API returns 200" +- ❌ NOT checking if code/implementation matches the spec + +**FOR requirements quality validation**: + +- ✅ "Are visual hierarchy requirements defined for all card types?" (completeness) +- ✅ "Is 'prominent display' quantified with specific sizing/positioning?" (clarity) +- ✅ "Are hover state requirements consistent across all interactive elements?" (consistency) +- ✅ "Are accessibility requirements defined for keyboard navigation?" (coverage) +- ✅ "Does the spec define what happens when logo image fails to load?" (edge cases) + +**Metaphor**: If your spec is code written in English, the checklist is its unit test suite. You're testing whether the requirements are well-written, complete, unambiguous, and ready for implementation - NOT whether the implementation works. + +## User Input + +```text +$ARGUMENTS +``` + +You **MUST** consider the user input before proceeding (if not empty). + +## Execution Steps + +1. **Setup**: Run `.specify/scripts/powershell/check-prerequisites.ps1 -Json` from repo root and parse JSON for FEATURE_DIR and AVAILABLE_DOCS list. + - All file paths must be absolute. + - For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot"). + +2. **Clarify intent (dynamic)**: Derive up to THREE initial contextual clarifying questions (no pre-baked catalog). They MUST: + - Be generated from the user's phrasing + extracted signals from spec/plan/tasks + - Only ask about information that materially changes checklist content + - Be skipped individually if already unambiguous in `$ARGUMENTS` + - Prefer precision over breadth + + Generation algorithm: + 1. Extract signals: feature domain keywords (e.g., auth, latency, UX, API), risk indicators ("critical", "must", "compliance"), stakeholder hints ("QA", "review", "security team"), and explicit deliverables ("a11y", "rollback", "contracts"). + 2. Cluster signals into candidate focus areas (max 4) ranked by relevance. + 3. Identify probable audience & timing (author, reviewer, QA, release) if not explicit. + 4. Detect missing dimensions: scope breadth, depth/rigor, risk emphasis, exclusion boundaries, measurable acceptance criteria. + 5. Formulate questions chosen from these archetypes: + - Scope refinement (e.g., "Should this include integration touchpoints with X and Y or stay limited to local module correctness?") + - Risk prioritization (e.g., "Which of these potential risk areas should receive mandatory gating checks?") + - Depth calibration (e.g., "Is this a lightweight pre-commit sanity list or a formal release gate?") + - Audience framing (e.g., "Will this be used by the author only or peers during PR review?") + - Boundary exclusion (e.g., "Should we explicitly exclude performance tuning items this round?") + - Scenario class gap (e.g., "No recovery flows detected—are rollback / partial failure paths in scope?") + + Question formatting rules: + - If presenting options, generate a compact table with columns: Option | Candidate | Why It Matters + - Limit to A–E options maximum; omit table if a free-form answer is clearer + - Never ask the user to restate what they already said + - Avoid speculative categories (no hallucination). If uncertain, ask explicitly: "Confirm whether X belongs in scope." + + Defaults when interaction impossible: + - Depth: Standard + - Audience: Reviewer (PR) if code-related; Author otherwise + - Focus: Top 2 relevance clusters + + Output the questions (label Q1/Q2/Q3). After answers: if ≥2 scenario classes (Alternate / Exception / Recovery / Non-Functional domain) remain unclear, you MAY ask up to TWO more targeted follow‑ups (Q4/Q5) with a one-line justification each (e.g., "Unresolved recovery path risk"). Do not exceed five total questions. Skip escalation if user explicitly declines more. + +3. **Understand user request**: Combine `$ARGUMENTS` + clarifying answers: + - Derive checklist theme (e.g., security, review, deploy, ux) + - Consolidate explicit must-have items mentioned by user + - Map focus selections to category scaffolding + - Infer any missing context from spec/plan/tasks (do NOT hallucinate) + +4. **Load feature context**: Read from FEATURE_DIR: + - spec.md: Feature requirements and scope + - plan.md (if exists): Technical details, dependencies + - tasks.md (if exists): Implementation tasks + + **Context Loading Strategy**: + - Load only necessary portions relevant to active focus areas (avoid full-file dumping) + - Prefer summarizing long sections into concise scenario/requirement bullets + - Use progressive disclosure: add follow-on retrieval only if gaps detected + - If source docs are large, generate interim summary items instead of embedding raw text + +5. **Generate checklist** - Create "Unit Tests for Requirements": + - Create `FEATURE_DIR/checklists/` directory if it doesn't exist + - Generate unique checklist filename: + - Use short, descriptive name based on domain (e.g., `ux.md`, `api.md`, `security.md`) + - Format: `[domain].md` + - If file exists, append to existing file + - Number items sequentially starting from CHK001 + - Each `/speckit.checklist` run creates a NEW file (never overwrites existing checklists) + + **CORE PRINCIPLE - Test the Requirements, Not the Implementation**: + Every checklist item MUST evaluate the REQUIREMENTS THEMSELVES for: + - **Completeness**: Are all necessary requirements present? + - **Clarity**: Are requirements unambiguous and specific? + - **Consistency**: Do requirements align with each other? + - **Measurability**: Can requirements be objectively verified? + - **Coverage**: Are all scenarios/edge cases addressed? + + **Category Structure** - Group items by requirement quality dimensions: + - **Requirement Completeness** (Are all necessary requirements documented?) + - **Requirement Clarity** (Are requirements specific and unambiguous?) + - **Requirement Consistency** (Do requirements align without conflicts?) + - **Acceptance Criteria Quality** (Are success criteria measurable?) + - **Scenario Coverage** (Are all flows/cases addressed?) + - **Edge Case Coverage** (Are boundary conditions defined?) + - **Non-Functional Requirements** (Performance, Security, Accessibility, etc. - are they specified?) + - **Dependencies & Assumptions** (Are they documented and validated?) + - **Ambiguities & Conflicts** (What needs clarification?) + + **HOW TO WRITE CHECKLIST ITEMS - "Unit Tests for English"**: + + ❌ **WRONG** (Testing implementation): + - "Verify landing page displays 3 episode cards" + - "Test hover states work on desktop" + - "Confirm logo click navigates home" + + ✅ **CORRECT** (Testing requirements quality): + - "Are the exact number and layout of featured episodes specified?" [Completeness] + - "Is 'prominent display' quantified with specific sizing/positioning?" [Clarity] + - "Are hover state requirements consistent across all interactive elements?" [Consistency] + - "Are keyboard navigation requirements defined for all interactive UI?" [Coverage] + - "Is the fallback behavior specified when logo image fails to load?" [Edge Cases] + - "Are loading states defined for asynchronous episode data?" [Completeness] + - "Does the spec define visual hierarchy for competing UI elements?" [Clarity] + + **ITEM STRUCTURE**: + Each item should follow this pattern: + - Question format asking about requirement quality + - Focus on what's WRITTEN (or not written) in the spec/plan + - Include quality dimension in brackets [Completeness/Clarity/Consistency/etc.] + - Reference spec section `[Spec §X.Y]` when checking existing requirements + - Use `[Gap]` marker when checking for missing requirements + + **EXAMPLES BY QUALITY DIMENSION**: + + Completeness: + - "Are error handling requirements defined for all API failure modes? [Gap]" + - "Are accessibility requirements specified for all interactive elements? [Completeness]" + - "Are mobile breakpoint requirements defined for responsive layouts? [Gap]" + + Clarity: + - "Is 'fast loading' quantified with specific timing thresholds? [Clarity, Spec §NFR-2]" + - "Are 'related episodes' selection criteria explicitly defined? [Clarity, Spec §FR-5]" + - "Is 'prominent' defined with measurable visual properties? [Ambiguity, Spec §FR-4]" + + Consistency: + - "Do navigation requirements align across all pages? [Consistency, Spec §FR-10]" + - "Are card component requirements consistent between landing and detail pages? [Consistency]" + + Coverage: + - "Are requirements defined for zero-state scenarios (no episodes)? [Coverage, Edge Case]" + - "Are concurrent user interaction scenarios addressed? [Coverage, Gap]" + - "Are requirements specified for partial data loading failures? [Coverage, Exception Flow]" + + Measurability: + - "Are visual hierarchy requirements measurable/testable? [Acceptance Criteria, Spec §FR-1]" + - "Can 'balanced visual weight' be objectively verified? [Measurability, Spec §FR-2]" + + **Scenario Classification & Coverage** (Requirements Quality Focus): + - Check if requirements exist for: Primary, Alternate, Exception/Error, Recovery, Non-Functional scenarios + - For each scenario class, ask: "Are [scenario type] requirements complete, clear, and consistent?" + - If scenario class missing: "Are [scenario type] requirements intentionally excluded or missing? [Gap]" + - Include resilience/rollback when state mutation occurs: "Are rollback requirements defined for migration failures? [Gap]" + + **Traceability Requirements**: + - MINIMUM: ≥80% of items MUST include at least one traceability reference + - Each item should reference: spec section `[Spec §X.Y]`, or use markers: `[Gap]`, `[Ambiguity]`, `[Conflict]`, `[Assumption]` + - If no ID system exists: "Is a requirement & acceptance criteria ID scheme established? [Traceability]" + + **Surface & Resolve Issues** (Requirements Quality Problems): + Ask questions about the requirements themselves: + - Ambiguities: "Is the term 'fast' quantified with specific metrics? [Ambiguity, Spec §NFR-1]" + - Conflicts: "Do navigation requirements conflict between §FR-10 and §FR-10a? [Conflict]" + - Assumptions: "Is the assumption of 'always available podcast API' validated? [Assumption]" + - Dependencies: "Are external podcast API requirements documented? [Dependency, Gap]" + - Missing definitions: "Is 'visual hierarchy' defined with measurable criteria? [Gap]" + + **Content Consolidation**: + - Soft cap: If raw candidate items > 40, prioritize by risk/impact + - Merge near-duplicates checking the same requirement aspect + - If >5 low-impact edge cases, create one item: "Are edge cases X, Y, Z addressed in requirements? [Coverage]" + + **🚫 ABSOLUTELY PROHIBITED** - These make it an implementation test, not a requirements test: + - ❌ Any item starting with "Verify", "Test", "Confirm", "Check" + implementation behavior + - ❌ References to code execution, user actions, system behavior + - ❌ "Displays correctly", "works properly", "functions as expected" + - ❌ "Click", "navigate", "render", "load", "execute" + - ❌ Test cases, test plans, QA procedures + - ❌ Implementation details (frameworks, APIs, algorithms) + + **✅ REQUIRED PATTERNS** - These test requirements quality: + - ✅ "Are [requirement type] defined/specified/documented for [scenario]?" + - ✅ "Is [vague term] quantified/clarified with specific criteria?" + - ✅ "Are requirements consistent between [section A] and [section B]?" + - ✅ "Can [requirement] be objectively measured/verified?" + - ✅ "Are [edge cases/scenarios] addressed in requirements?" + - ✅ "Does the spec define [missing aspect]?" + +6. **Structure Reference**: Generate the checklist following the canonical template in `.specify/templates/checklist-template.md` for title, meta section, category headings, and ID formatting. If template is unavailable, use: H1 title, purpose/created meta lines, `##` category sections containing `- [ ] CHK### ` lines with globally incrementing IDs starting at CHK001. + +7. **Report**: Output full path to created checklist, item count, and remind user that each run creates a new file. Summarize: + - Focus areas selected + - Depth level + - Actor/timing + - Any explicit user-specified must-have items incorporated + +**Important**: Each `/speckit.checklist` command invocation creates a checklist file using short, descriptive names unless file already exists. This allows: + +- Multiple checklists of different types (e.g., `ux.md`, `test.md`, `security.md`) +- Simple, memorable filenames that indicate checklist purpose +- Easy identification and navigation in the `checklists/` folder + +To avoid clutter, use descriptive types and clean up obsolete checklists when done. + +## Example Checklist Types & Sample Items + +**UX Requirements Quality:** `ux.md` + +Sample items (testing the requirements, NOT the implementation): + +- "Are visual hierarchy requirements defined with measurable criteria? [Clarity, Spec §FR-1]" +- "Is the number and positioning of UI elements explicitly specified? [Completeness, Spec §FR-1]" +- "Are interaction state requirements (hover, focus, active) consistently defined? [Consistency]" +- "Are accessibility requirements specified for all interactive elements? [Coverage, Gap]" +- "Is fallback behavior defined when images fail to load? [Edge Case, Gap]" +- "Can 'prominent display' be objectively measured? [Measurability, Spec §FR-4]" + +**API Requirements Quality:** `api.md` + +Sample items: + +- "Are error response formats specified for all failure scenarios? [Completeness]" +- "Are rate limiting requirements quantified with specific thresholds? [Clarity]" +- "Are authentication requirements consistent across all endpoints? [Consistency]" +- "Are retry/timeout requirements defined for external dependencies? [Coverage, Gap]" +- "Is versioning strategy documented in requirements? [Gap]" + +**Performance Requirements Quality:** `performance.md` + +Sample items: + +- "Are performance requirements quantified with specific metrics? [Clarity]" +- "Are performance targets defined for all critical user journeys? [Coverage]" +- "Are performance requirements under different load conditions specified? [Completeness]" +- "Can performance requirements be objectively measured? [Measurability]" +- "Are degradation requirements defined for high-load scenarios? [Edge Case, Gap]" + +**Security Requirements Quality:** `security.md` + +Sample items: + +- "Are authentication requirements specified for all protected resources? [Coverage]" +- "Are data protection requirements defined for sensitive information? [Completeness]" +- "Is the threat model documented and requirements aligned to it? [Traceability]" +- "Are security requirements consistent with compliance obligations? [Consistency]" +- "Are security failure/breach response requirements defined? [Gap, Exception Flow]" + +## Anti-Examples: What NOT To Do + +**❌ WRONG - These test implementation, not requirements:** + +```markdown +- [ ] CHK001 - Verify landing page displays 3 episode cards [Spec §FR-001] +- [ ] CHK002 - Test hover states work correctly on desktop [Spec §FR-003] +- [ ] CHK003 - Confirm logo click navigates to home page [Spec §FR-010] +- [ ] CHK004 - Check that related episodes section shows 3-5 items [Spec §FR-005] +``` + +**✅ CORRECT - These test requirements quality:** + +```markdown +- [ ] CHK001 - Are the number and layout of featured episodes explicitly specified? [Completeness, Spec §FR-001] +- [ ] CHK002 - Are hover state requirements consistently defined for all interactive elements? [Consistency, Spec §FR-003] +- [ ] CHK003 - Are navigation requirements clear for all clickable brand elements? [Clarity, Spec §FR-010] +- [ ] CHK004 - Is the selection criteria for related episodes documented? [Gap, Spec §FR-005] +- [ ] CHK005 - Are loading state requirements defined for asynchronous episode data? [Gap] +- [ ] CHK006 - Can "visual hierarchy" requirements be objectively measured? [Measurability, Spec §FR-001] +``` + +**Key Differences:** + +- Wrong: Tests if the system works correctly +- Correct: Tests if the requirements are written correctly +- Wrong: Verification of behavior +- Correct: Validation of requirement quality +- Wrong: "Does it do X?" +- Correct: "Is X clearly specified?" diff --git a/.claude/commands/speckit.clarify.md b/.claude/commands/speckit.clarify.md new file mode 100644 index 0000000..0678e92 --- /dev/null +++ b/.claude/commands/speckit.clarify.md @@ -0,0 +1,181 @@ +--- +description: Identify underspecified areas in the current feature spec by asking up to 5 highly targeted clarification questions and encoding answers back into the spec. +handoffs: + - label: Build Technical Plan + agent: speckit.plan + prompt: Create a plan for the spec. I am building with... +--- + +## User Input + +```text +$ARGUMENTS +``` + +You **MUST** consider the user input before proceeding (if not empty). + +## Outline + +Goal: Detect and reduce ambiguity or missing decision points in the active feature specification and record the clarifications directly in the spec file. + +Note: This clarification workflow is expected to run (and be completed) BEFORE invoking `/speckit.plan`. If the user explicitly states they are skipping clarification (e.g., exploratory spike), you may proceed, but must warn that downstream rework risk increases. + +Execution steps: + +1. Run `.specify/scripts/powershell/check-prerequisites.ps1 -Json -PathsOnly` from repo root **once** (combined `--json --paths-only` mode / `-Json -PathsOnly`). Parse minimal JSON payload fields: + - `FEATURE_DIR` + - `FEATURE_SPEC` + - (Optionally capture `IMPL_PLAN`, `TASKS` for future chained flows.) + - If JSON parsing fails, abort and instruct user to re-run `/speckit.specify` or verify feature branch environment. + - For single quotes in args like "I'm Groot", use escape syntax: e.g 'I'\''m Groot' (or double-quote if possible: "I'm Groot"). + +2. Load the current spec file. Perform a structured ambiguity & coverage scan using this taxonomy. For each category, mark status: Clear / Partial / Missing. Produce an internal coverage map used for prioritization (do not output raw map unless no questions will be asked). + + Functional Scope & Behavior: + - Core user goals & success criteria + - Explicit out-of-scope declarations + - User roles / personas differentiation + + Domain & Data Model: + - Entities, attributes, relationships + - Identity & uniqueness rules + - Lifecycle/state transitions + - Data volume / scale assumptions + + Interaction & UX Flow: + - Critical user journeys / sequences + - Error/empty/loading states + - Accessibility or localization notes + + Non-Functional Quality Attributes: + - Performance (latency, throughput targets) + - Scalability (horizontal/vertical, limits) + - Reliability & availability (uptime, recovery expectations) + - Observability (logging, metrics, tracing signals) + - Security & privacy (authN/Z, data protection, threat assumptions) + - Compliance / regulatory constraints (if any) + + Integration & External Dependencies: + - External services/APIs and failure modes + - Data import/export formats + - Protocol/versioning assumptions + + Edge Cases & Failure Handling: + - Negative scenarios + - Rate limiting / throttling + - Conflict resolution (e.g., concurrent edits) + + Constraints & Tradeoffs: + - Technical constraints (language, storage, hosting) + - Explicit tradeoffs or rejected alternatives + + Terminology & Consistency: + - Canonical glossary terms + - Avoided synonyms / deprecated terms + + Completion Signals: + - Acceptance criteria testability + - Measurable Definition of Done style indicators + + Misc / Placeholders: + - TODO markers / unresolved decisions + - Ambiguous adjectives ("robust", "intuitive") lacking quantification + + For each category with Partial or Missing status, add a candidate question opportunity unless: + - Clarification would not materially change implementation or validation strategy + - Information is better deferred to planning phase (note internally) + +3. Generate (internally) a prioritized queue of candidate clarification questions (maximum 5). Do NOT output them all at once. Apply these constraints: + - Maximum of 10 total questions across the whole session. + - Each question must be answerable with EITHER: + - A short multiple‑choice selection (2–5 distinct, mutually exclusive options), OR + - A one-word / short‑phrase answer (explicitly constrain: "Answer in <=5 words"). + - Only include questions whose answers materially impact architecture, data modeling, task decomposition, test design, UX behavior, operational readiness, or compliance validation. + - Ensure category coverage balance: attempt to cover the highest impact unresolved categories first; avoid asking two low-impact questions when a single high-impact area (e.g., security posture) is unresolved. + - Exclude questions already answered, trivial stylistic preferences, or plan-level execution details (unless blocking correctness). + - Favor clarifications that reduce downstream rework risk or prevent misaligned acceptance tests. + - If more than 5 categories remain unresolved, select the top 5 by (Impact * Uncertainty) heuristic. + +4. Sequential questioning loop (interactive): + - Present EXACTLY ONE question at a time. + - For multiple‑choice questions: + - **Analyze all options** and determine the **most suitable option** based on: + - Best practices for the project type + - Common patterns in similar implementations + - Risk reduction (security, performance, maintainability) + - Alignment with any explicit project goals or constraints visible in the spec + - Present your **recommended option prominently** at the top with clear reasoning (1-2 sentences explaining why this is the best choice). + - Format as: `**Recommended:** Option [X] - ` + - Then render all options as a Markdown table: + + | Option | Description | + |--------|-------------| + | A |