Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

statement-normalizer

CILicense: Apache-2.0

Turn messy bank and credit-card statements into clean, normalized transaction JSON. Deterministic, rule-based, dependency-free, and easy to audit.

statement-normalizer reads the statement formats banks and cards actually export, CSV, OFX/QFX, MT940, CAMT.053, QIF, and line-oriented text, and emits a single normalized transaction schema with a consistent sign convention, parsed dates, and proper decimal amounts. Given the same input it produces the same output every time, which is exactly what you want when the data feeds bookkeeping, reconciliation, or any financial workflow.

Why

Every bank exports statements differently: some give you a single signed Amount column, others split Debit / Credit, some wrap negatives in parentheses, some hand you OFX/QFX or QIF, some only give you a PDF. Downstream tools all want the same thing: a tidy list of transactions with a consistent sign convention, parsed dates, and proper decimal amounts.

statement-normalizer does exactly that, and does it well. It is a focused, well-tested building block you can drop into an import pipeline: it parses the real-world export shapes, normalizes the sign convention, and de-duplicates across overlapping statements, so the rest of your pipeline only ever sees one clean schema.

What it does

  • Parses the statement shapes banks and cards actually export:
    • CSV — signed-amount or separate debit/credit columns, optional running balance, fuzzy header detection with a broad alias table covering the real-world headers from Chase, Bank of America, Wells Fargo, Amex, Capital One, Discover and many others.
    • OFX / QFX — both the SGML (OFX 1.x) and XML (OFX 2.x) export styles, using a small built-in tokenizer (no external OFX dependency).
    • MT940 — the SWIFT bank-statement message format (.sta / .940), common for European business accounts.
    • CAMT.053 — ISO 20022 BankToCustomerStatement XML, the modern standard replacing MT940. Parsed with the standard-library XML parser, namespace-agnostic.
    • CAMT.052 — ISO 20022 BankToCustomerAccountReport XML, the intra-day / interim report banks send before the end-of-day CAMT.053 is cut. Same entry model as CAMT.053; auto-detected from the container element so the two XML shapes never get confused.
    • QIF — Quicken Interchange Format (.qif), the long-lived plain-text export used by Quicken and many banks and personal-finance tools. Parsed with a small built-in tokenizer (no external QIF dependency).
    • Text — line-oriented statement text, e.g. the output of running a PDF through a text extractor like pdftotext.

Format support matrix

FormatExtensionsDirection sourceAccount idStrong dedup keyMulti-currency
CSV (signed).csvsign of Amount (or --invert-amounts)content hashper-row Currency col
CSV (debit/credit).csvwhich column is populatedcontent hashper-row Currency col
OFX / QFX.ofx.qfxsigned TRNAMTACCTIDFITIDCURDEF
MT940.sta.940.mt940D / C mark on :61::25:content hash:60F: currency
CAMT.053.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
CAMT.052.xml (sniffed)CdtDbtInd (DBIT/CRDT)IBAN / Othr idAcctSvcrRefAmt@Ccy
QIF.qifsign of T / U amountcontent hash (N ref hint)--currency default
Text / PDF-text.txtsign / column heuristiccontent hash--currency default

All amounts are normalized to one sign convention (debits negative, credits positive) regardless of how the source expressed direction.

  • Emits a single normalized Transaction schema with a consistent sign convention (debits negative, credits positive), parsed date, Decimalamount, derived txn_type, optional balance, fitid, account_id, and currency.
  • De-duplicates transactions with deterministic heuristics:
    • Bank-assigned FITID (from OFX) is the authoritative key when present.
    • Otherwise a content hash over (date, signed amount, currency, canonical description), while preserving legitimately-repeated same-day charges.
  • Merges multiple statements (e.g. overlapping monthly exports) into one de-duplicated list.
  • Ships a CLI and a small Python API.

Boundaries

  • Binary PDF workflows. Pipe your PDF through a text extractor first, then feed the text in.
  • Transaction categorization. Pair normalized output with your preferred categorization layer.

Install

pip install statement-normalizer

Or from a checkout:

pip install -e .

Python 3.9+. No runtime dependencies.

Usage — CLI

# Normalize a single file (format auto-detected by extension/content)
statement-normalizer statement.csv --pretty
# Force a format
statement-normalizer --format ofx export.qfx
# MT940, CAMT.053, and QIF work the same way
statement-normalizer export.sta --stats
statement-normalizer --format camt053 statement.xml
statement-normalizer export.qif --pretty
# Merge overlapping monthly exports, de-duplicating across them
statement-normalizer jan.csv feb.csv --merge -o all.json
# Flat CSV instead of JSON
statement-normalizer statement.ofx --csv -o transactions.csv
# Print a summary (counts, totals, date range) to stderr while still emitting JSON
statement-normalizer statement.csv --stats
# Read from stdin ('-')
cat statement.csv | statement-normalizer --format csv -
# Issuers that report charges as positive / payments as negative (e.g. some# credit-card exports) — flip the sign into the debit-negative convention
statement-normalizer amex.csv --invert-amounts
# Set a default currency when the source omits one
statement-normalizer eu_statement.csv --currency EUR

CLI flags

FlagEffect
--format {csv,ofx,text,mt940,camt053,qif}force input format (default: auto-detect)
--csvemit a flat transactions CSV instead of JSON
--statsprint a counts/totals/date-range summary to stderr
--mergemerge all inputs into one cross-statement-deduped list
--no-dedupdisable de-duplication
--invert-amountsflip the sign of CSV single-Amount-column values
--currency CCYdefault currency when the source omits one
--prettypretty-print JSON
-o FILEwrite to a file instead of stdout

Usage — Python

fromstatement_normalizerimportnormalize_filestatement=normalize_file("statement.csv")
print(statement.account_id, statement.currency)
fortxninstatement.transactions:
print(txn.date, txn.amount, txn.txn_type.value, txn.description)
# JSON-serializable dictimportjsonprint(json.dumps(statement.to_dict(), indent=2))

Merging multiple files with cross-statement dedup:

fromstatement_normalizer.normalizeimportnormalize_manytxns=normalize_many(["jan.ofx", "feb.ofx"])

Examples

The examples/ directory ships synthetic statements in the shapes real banks and cards export, plus a runnable tour:

python examples/demo.py
FileShape it demonstrates
chase_checking.csvChase Details/Posting Date/Description/Amount/Type/Balance
bofa_checking.csvBank of America Date/Description/Amount/Running Bal.
wells_fargo_checking.csvWells Fargo Date/Amount/*/Payee/Memo (blank Payee, text in Memo)
amex_creditcard.csvAmex single Amount with charges positive (use --invert-amounts)
capital_one_creditcard.csvCapital One split Debit/Credit columns
discover_creditcard.csvDiscover Trans. Date/Amount/Category (charges positive)
sample.mt940SWIFT MT940 (:25:/:60F:/:61:/:86:)
sample.camt053.xmlISO 20022 CAMT.053 (<Ntry> / CdtDbtInd)
sample.qifQuicken QIF (!Type:Bank, D/T/P/M/N/L fields)
overlap_jan.csv, overlap_feb.csvtwo overlapping months for the dedup/merge demo

Dedup/merge across overlapping months in one line:

statement-normalizer examples/overlap_jan.csv examples/overlap_feb.csv --merge --stats
# 12 raw rows across the two files -> 9 after the 3 overlapping rows collapse,# while a legitimately-repeated same-merchant charge is preserved.

Example output

Input CSV:

Transaction Date,Description,Amount2024-02-01,GAS STATION 4471,(45.20)2024-02-05,PAYMENT - THANK YOU,250.00

Output JSON (abridged):

{
"account_id": null,
"currency": "USD",
"source_format": "csv",
"transaction_count": 2,
"transactions": [
{
"date": "2024-02-01",
"amount": "-45.20",
"description": "GAS STATION 4471",
"txn_type": "debit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
},
{
"date": "2024-02-05",
"amount": "250.00",
"description": "PAYMENT - THANK YOU",
"txn_type": "credit",
"currency": "USD",
"balance": null,
"fitid": null,
"account_id": null,
"source_format": "csv"
}
]
}

Sign convention

amount is always signed: negative = money out (debit), positive = money in (credit). txn_type is derived from the sign for convenience. Parsers normalize to this convention regardless of how the source expressed direction (separate debit columns, parentheses, trailing minus, etc.).

Development

pip install -e ".[dev]"
pytest
python examples/demo.py # runnable tour over the synthetic examples

The test suite runs entirely over synthetic sample statements committed under tests/fixtures/ and examples/. No real account data is used anywhere in this project; please keep it that way and only commit synthetic fixtures.

CI runs the tests on Python 3.9–3.13 (plus a macOS/Windows spot-check) and builds the sdist + wheel on every push and pull request.

Agent interface

For AI agents and automation: the CLI reads a statement on stdin (-) and prints normalized transaction JSON on stdout. Add --json-errors to get a structured error envelope on stderr ({"ok": false, "error": {"type": "parse_error", "message": ...}}). Exit codes are stable: 0 success, 1 parse/input error, 2 usage error. The full machine interface is in AGENT.md and llms.txt. This tool also backs the normalize_bank_statement tool in maxed-mcp.

License

Apache-2.0.

About

Deterministic bank and card statement parser that converts CSV, OFX/QFX, MT940, CAMT.053, and text exports into normalized transaction JSON.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages