Skip to content
Use this GitHub action with your project
Add this Action to an existing workflow or create a new one
View on Marketplace

Repository files navigation

Important

This repository has moved into the Governed Agent Stack monorepo.

Active development is now at packages/sql-sop/. This repo is archived and read-only. Its full commit history is preserved here; new work, issues and releases happen in the monorepo.


sql-sop

PyPIDownloads

CIPyPIPythonLicenseDiscordMarketplacePlaygroundDownloadscodecov

Part of the Governed Agent Stack: free, on-prem building blocks for an AI agent you can point at a real database and audit.

Links

Where this leads: sql-steward

sql-sop lints the SQL people write. When you want an AI agent to write and run SQL against a real database without ever handling a connection string, that is sql-steward, the governed MCP gateway from the same stack. The agent never writes raw SQL: every query is compiled from a semantic layer you control, blocked PII is refused before retrieval, and every call lands in a tamper-evident audit. sql-sop guards the SQL humans write; sql-steward guards the SQL an agent generates.

Companion tools

sql-sop also composes with another small on-prem tool:

  • sql-explorer-mcp: read-only MCP server that lets an AI introspect and query SQL Server / Postgres / SQLite; sql-sop is one of its safety layers, rejecting dangerous queries before they run.

The dbt-aware rule pack (DBT001+) extends sql-sop into dbt projects. See the ADR for the broader roadmap.

Why Does This Exist?

One bad SQL query can delete production data, expose customer records, or bring down a database. Most teams only find out after the damage is done. sql-sop catches dangerous patterns automatically, before the query ever runs, in 0.08 seconds.

Key Numbers

Rules48 (9 errors, 25 warnings, 3 structural, 6 T-SQL, 5 Python-source); 53 with --contract
Tests303
Scan speed0.08s across 200 files
PyPI installs2,000+ (mirrors excluded)
Version0.9.0

Fluent API (v0.2.0)

fromsql_guardimportSqlGuardresult=SqlGuard().enable("E001", "W001").scan("DELETE FROM users")
print(result.passed) # Falseprint(result.summary()) # "1 error, 0 warnings in 1 statement"

Fast, rule-based SQL linter. 48 rules (43 SQL + 5 Python), with an optional Contracts pack (5 schema-aware rules) when you supply --contract path.yml. SQL Server-focused rules for T-SQL shops. Inline disable, project config, git-changed-only mode, and SARIF output for GitHub Code Scanning. 2,000+ installs on PyPI.

Catches dangerous SQL before it reaches production -- DELETE without WHERE, UPDATE without WHERE, SQL injection patterns, SELECT *, contract drift, and 40+ more. Runs as a CLI tool, pre-commit hook, and GitHub Action.

Built to catch dangerous patterns like DELETE without WHERE before they ever reach a production database.

Want the same rules inside your AI assistant? sql-sop-mcp exposes them over MCP, so the model is told a query is unsafe before it suggests it.


Quick start

If sql-sop catches a real bug for you, a GitHub star is the easiest way to help. It makes the project more discoverable for people with the same problem.

pip install sql-sop
sql-sop check .# Also scan .py files for SQL hazards in execute()/read_sql() calls:
pip install "sql-sop[python]"
sql-sop check . --include-python
queries/create_orders.sql
L3: ERROR [E001] DELETE without WHERE clause -- this will delete all rows
-> Add a WHERE clause to limit affected rows
L7: WARN [W001] SELECT * -- specify columns explicitly
-> Replace with: SELECT col1, col2, col3 FROM ...
Found 2 issues (1 error, 1 warning) in 1 file (0.001s)

Where sql-sop runs

Most teams have no SQL review process at all. sql-sop gives you one in two places, both driven by the same rules:

  • Pre-commit -- runs in under 0.2s on the files you changed, so a bad DELETE FROM users is blocked before it is ever committed.
  • CI -- runs on every pull request and uploads SARIF, so findings appear inline in the GitHub "Files changed" view through Code Scanning.

Same engine in both, no AI, no API keys, nothing to host. Fast enough to sit in the commit path, strict enough to gate a PR.


Set up (5 minutes)

Step 1: Pre-commit hook

# .pre-commit-config.yamlrepos:
- repo: https://github.com/Pawansingh3889/sql-soprev: v0.9.0hooks:
- id: sql-sopargs: [--severity, error] # only block on errors locally
pip install pre-commit
pre-commit install

Now every git commit with .sql changes runs sql-sop automatically. Errors block the commit. Warnings are shown but don't block.

Step 2: GitHub Actions

# .github/workflows/sql-quality.ymlname: SQL Qualityon:
pull_request:
paths: ['**/*.sql']permissions:
contents: readpull-requests: writejobs:
lint:
runs-on: ubuntu-lateststeps:
- uses: actions/checkout@v4
- uses: Pawansingh3889/sql-sop@v1with:
severity: warning

That's it. Every SQL change gets an instant, rule-based lint on the PR.

Step 3 (optional): CLI for manual checks

pip install sql-sop
sql-sop check .# scan current directory
sql-sop check queries/ --severity error # errors only
sql-sop check . --fail-fast # stop on first error
sql-sop check . --disable E002 W008 # skip specific rules
sql-sop list-rules # show every registered rule

Rules

Errors (block commit by default)

IDNameWhat it catches
E001delete-without-whereDELETE FROM orders; -- deletes all rows
E002drop-without-if-existsDROP TABLE users; -- fails if table missing
E003grant-revokeGRANT SELECT ON users TO public; -- privilege escalation
E004string-concat-in-whereWHERE id = '' + @input -- SQL injection
E005insert-without-columnsINSERT INTO t VALUES (...) -- breaks on schema change
E006update-without-whereUPDATE orders SET status = 'x'; -- overwrites every row
E007alter-add-not-null-no-defaultALTER TABLE t ADD c INT NOT NULL; -- locks table for full rewrite
E008drop-columnALTER TABLE t DROP COLUMN c; -- irreversible, breaks subscribers

Warnings (advisory by default)

IDNameWhat it catches
W001select-starSELECT * FROM users -- pulls unnecessary columns
W002missing-limitUnbounded SELECT -- could return millions of rows
W003function-on-columnWHERE YEAR(date) = 2024 -- kills index usage
W004missing-aliasJOIN without table aliases -- hard to read
W005subquery-in-whereWHERE x IN (SELECT ...) -- often slower than JOIN
W006orderby-without-limitORDER BY without LIMIT -- sorts entire result
W007hardcoded-valuesWHERE amount > 10000 -- use parameters
W008mixed-case-keywordsselect ... FROM -- inconsistent casing
W009missing-semicolonStatement not terminated with ;
W010commented-out-code-- SELECT * FROM old_table -- use version control
W011union-without-allUNION between disjoint sets -- forces a deduplication sort, UNION ALL is faster when uniqueness is guaranteed
W012group-by-ordinalGROUP BY 1, 2 -- fragile to SELECT-list reorders
W013window-missing-partitionOVER () -- unpredictable results and unclear intent
W014case-without-elseCASE WHEN ... THEN ... END -- unmatched rows return NULL
W015join-function-on-columnJOIN customers c ON UPPER(o.email) = UPPER(c.email) -- kills index seek
W016not-in-with-subqueryWHERE id NOT IN (SELECT ...) -- silently returns 0 rows on NULL
W017leading-wildcard-likeWHERE name LIKE '%smith' -- non-SARGable, full scan
W018or-across-columnsWHERE a = 1 OR b = 2 -- defeats single-column indexes
W019count-distinct-unboundedCOUNT(DISTINCT col) with no WHERE / GROUP BY / LIMIT -- full sort + distinct over the whole table
W020truncate-tableTRUNCATE TABLE staging; -- bypasses triggers, resets identity
W022cross-join-explicitFROM products CROSS JOIN regions -- Cartesian product, confirm intent
W023scalar-udf-in-whereWHERE dbo.fn_X(col) = 1 -- row-by-row predicate evaluation
W024select-distinct-suspiciousSELECT DISTINCT a, b FROM x JOIN y ON ... -- DISTINCT often masks a missing join condition or GROUP BY
W025assertion-malformed-- @assert: <predicate> comment whose predicate does not match the sql-sop grammar (row_count <op> <int>, unique(<col>), not_null(<col>), <col> <op> <literal>)

Structural (v0.3.0+, sqlparse-based)

Rules that need an AST view of the statement, parsed via sqlparse. Catch issues that single-line regex matching cannot reliably see.

IDNameWhat it catches
S001implicit-cross-joinJOIN customers with no ON / USING -- accidental Cartesian product
S002deeply-nested-subquerySubqueries beyond 3 levels deep -- typically a refactor opportunity
S003unused-cteWITH x AS (...) defined but never referenced

T-SQL (v0.5.0+)

Rules targeting SQL Server anti-patterns common in legacy stored procs and SSRS datasets. Fire on text patterns that do not appear in BigQuery or Postgres code, so they run unconditionally with near-zero false positives on non-T-SQL input.

IDNameWhat it catches
T001with-nolockSELECT * FROM t WITH (NOLOCK) -- dirty reads
T002xp-cmdshellEXEC xp_cmdshell ... -- shell-exec surface
T003cursor-declarationDECLARE c CURSOR FOR ... -- row-by-row processing
T004deprecated-outer-joinWHERE a.x *= b.y -- removed in SQL Server 2012+
T005create-index-without-onlineCREATE INDEX ix ON t (...) -- locks table; add WITH (ONLINE = ON)
T006select-into-without-typed-fieldsSELECT * INTO target FROM source -- destination schema is inferred at runtime

Contracts (opt-in via --contract)

Pass --contract path/to/contract.yml (or set contract: in .sql-guard.yml) to lint queries against the expected schema. Without a contract these rules are silent. Format is a thin subset of the open data-contract space; see tests/fixtures/contract_sample.yml for a working example.

IDNameWhat it catches
C001column-not-in-contractSELECT o.bogus FROM orders o -- column not declared for that table
C002table-not-in-contractSELECT * FROM ghost_table -- table absent from the contract
C003not-null-violationINSERT INTO orders (id) VALUES (1) -- omits a NOT NULL column
C004primary-key-missing-on-insertINSERT omits a PK column with no default
C005unmapped-fkJOIN ... ON o.id = c.id -- columns have no FK relationship in the contract

Two helper subcommands round out the workflow:

# Bootstrap a contract from an existing database (requires sql-sop[snapshot]):
sql-sop schema-snapshot \
--dsn "mssql+pyodbc://user:pass@host/db?driver=ODBC+Driver+18+for+SQL+Server" \
--output contract.yml
# Validate a contract YAML structure before running rules (CI-friendly):
sql-sop validate-contract --contract contract.yml

Python scanning (v0.4.0+, opt-in)

Enable with pip install "sql-sop[python]" and --include-python. Uses libCST to walk Python source and extract SQL strings from .execute(), .read_sql(), sqlalchemy.text(...) calls and sql =/query = style assignments. Then applies every rule above, plus four that only make sense at the Python level:

IDNameWhat it catches
P001fstring-in-executecursor.execute(f"... {user_input}") -- SQL injection
P002concat-in-executecursor.execute("..." + user_input) -- SQL injection
P003format-in-execute.format() or % interpolation into an execute call
P004bare-variable-in-executecursor.execute(query) where query is an unchecked variable
P005sqlalchemy-text-fstringsqlalchemy.text(f"... {var}") -- SQL injection on the SQLAlchemy text() surface

Configuration

Disable specific rules

sql-sop check . --disable E002 W008 W010

Project config file (.sql-guard.yml)

Drop a .sql-guard.yml (or .sql-guard.yaml) at the repo root. The loader walks up from the current directory; CLI flags merge with and override these settings.

disable:
- W005
- T001ignore:
- migrations/legacy/
- vendor/include_python: trueseverity: warning

Inline disable comments

Silence a known false positive on a single line, no project-wide override needed:

SELECT*FROM lookups; -- sql-guard: disable=W001SELECT*FROM users -- sql-guard: disable=W001,W002WHERE name LIKE'%smith';
-- sql-guard: disable-next-line=W017SELECT*FROM events WHERE name LIKE'%checkout';

A bare -- sql-guard: disable (no equals sign) silences every rule on the line. The same directives work in Python with # instead of --.

Lint only changed files

For pre-commit and CI on big repos:

sql-sop check . --changed-only # working tree
sql-sop check . --changed-only --changed-base main # vs a branch ref

Falls back to a full scan with a warning when not in a git repo.

Severity filtering

sql-sop check . --severity error # only show errors
sql-sop check . --severity warning # show everything (default)

Fail fast

sql-sop check . --fail-fast # stop after first error found

SARIF output for GitHub Code Scanning

Render findings inline on PRs in the GitHub Files Changed view:

sql-sop check . --format sarif --output results.sarif

In a GitHub Actions workflow:

- run: sql-sop check . --format sarif --output sql-guard.sarif
- uses: github/codeql-action/upload-sarif@v3with:
sarif_file: sql-guard.sarif

Performance

sql-guard is designed to be fast:

  • Compiled regex -- patterns compiled once at startup, reused per file
  • Two-pass scanning -- single-line rules run first (8 of 20 SQL rules), multi-line parsing only when needed
  • Line-by-line streaming -- files read line by line, not loaded entirely into memory
  • Early exit -- --fail-fast stops on first error
Benchmark: 200 SQL files, 20 SQL rules
sql-guard: 0.08 seconds
sqlfluff: 45 seconds (560x slower)

Production Use Case

In a regulated data environment, sql-sop runs as a pre-commit hook on all SQL that touches the database. Combined with read-only database users and container isolation, it forms part of a layered safety setup that prevents accidental writes to production.


How it compares

sql-sopsqlfluffsql-lint
Rules48 (focused)800+ (comprehensive)~20
Speed<0.1s for 200 files45s for 200 files~2s
Config neededZeroExtensiveMinimal
LanguagePythonPythonJavaScript
Pre-commitYesYesNo
GitHub ActionYesCommunityNo
AI integrationMCP server (sql-sop-mcp)NoNo

sql-sop is not a replacement for sqlfluff. It's a fast first pass that catches 80% of real issues with zero setup. If you need dialect-specific formatting and 800 rules, use sqlfluff. If you want instant feedback on dangerous SQL, use sql-guard.


Contributing

git clone https://github.com/Pawansingh3889/sql-guard.git
cd sql-guard
pip install -e ".[dev]"
pytest

Adding a new rule

  1. Create a class in sql_guard/rules/errors.py or warnings.py
  2. Inherit from Rule, set id, name, severity, description
  3. Override check_line() for single-line rules or check_statement() for multi-line
  4. Add to ALL_RULES in sql_guard/rules/__init__.py
  5. Add a test in tests/test_rules.py
  6. Add a trigger case in tests/fixtures/
classMyNewRule(Rule):
id="W011"name="my-rule"severity="warning"description="What this rule catches"multiline=False_pattern=Rule._compile(r"your regex here")
defcheck_line(self, line, line_number, file):
ifself._pattern.search(line):
returnFinding(
rule_id=self.id,
severity=self.severity,
file=file,
line=line_number,
message="What went wrong",
suggestion="How to fix it",
)
returnNone

PRs welcome. Keep rules simple, keep patterns fast.


Contributors

Thank you to the people who have shipped rules and code to sql-sop.

ContributorContribution
@tmchowW011 union-without-all. Flags UNION where UNION ALL would be safe and faster.
@tmchowP005 sqlalchemy-text-fstring. Catches sqlalchemy.text(f"...{var}") patterns that defeat parameter binding.
@mvanhornW019 count-distinct-unbounded. Flags COUNT(DISTINCT col) without WHERE, GROUP BY, or LIMIT.
@mvanhornW015 join-function-on-column. JOIN-side companion to W003. Flags function calls wrapping columns inside JOIN ... ON predicates.
@mvanhornW023 scalar-udf-in-where. Flags schema-qualified scalar UDF calls inside WHERE, HAVING, and ON predicates.
@Prabhu-1409W013 window-without-partition. Flags OVER () without PARTITION BY, dialect-aware messaging for Postgres and Redshift.
@hellozzmW014 case-without-else. Walks CASE/END token-by-token; catches outer CASE without ELSE even when an inner CASE does have one.
@vibeyclawW022 cross-join-explicit. Flags explicit CROSS JOIN. Strips trailing line comments before matching to avoid false positives on commentary.

See the full contributors graph on GitHub.

Want to add your name here? Pick a good first issue, follow CONTRIBUTING.md, and check the roadmap for the next batch of rules. v0.7 just shipped (contracts pack); v0.8 is shaping up around dialect-aware coverage.


License

MIT

About

Fast rule-based SQL linter on PyPI (sql-sop). 48 rules (T-SQL, Python-source, contracts), 303 tests, SARIF output, pre-commit hook + GitHub Action. On-prem, zero-config.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages