Skip to content

Repository files navigation

PaperlessLabelAgent

The paperless label agent classifies documents for Paperless-ngx, proposing tags, a correspondent and a document type for each file in a dedicated folder. It asks you to confirm every proposition before finalizing the classification.

It reads the tags, correspondents and document types that already exist in your Paperless-ngx instance. Then, the Agent uses a local LLM (via Ollama) to either match a document against those existing entities or, if no enity fits, propose new ones. You review every proposal for its correctness. Any rejected entity (either existing or new one) is reclassified automatically, with your rejection fed back into the next attempt so the model doesn't repeat the same mistake. Repetion is constrained by a threshold value.

For this, two execution strategies are available (see STRATEGY below):

  • iterative (default) — classifies, reviews and confirms one document fully before moving to the next. This strategy is recommended because it avoids the need for unifying all newly recommended entities based on their semantics.
  • sequential (default) — classifies every document first, then reviews them one by one. Requires functionality for unifying all newly recommended entities based on their semantics (not yet implemented)

Requirements

  • Python 3.12+
  • Ollama, running locally with a model pulled that supports structured/tool output
  • Tesseract OCR installed locally
  • A running Paperless-ngx instance — or use mock mode (see below) to try the agent without one

Configuration

Create a .env file in the project root with the following keys:

KeyDescription
API_URLBase URL of your Paperless-ngx API, e.g. http://localhost:8000/api
ACCOUNTPaperless-ngx username
PASSWORDPaperless-ngx password
MODELOllama model tag to use for classification, e.g. qwen3.5:9b or llama3:8b
TESSDATA_PATHPath to your Tesseract tessdata directory
OCR_LANGUAGESLanguages that should be handled by Tesseract OCR
INPUT_FOLDERFolder containing the PDF documents to classify
STRATEGYExecution strategy: sequential or iterative (default)
ENTITY_LANGUAGELanguage(s) the LLMs are promted to provide entities for

Mock mode: if ACCOUNT and PASSWORD both contain the string mock, the agent fetches sample tags/correspondents/document types from test/paperless-instance-mock instead of using the Paperless-ngx API on an running instance. The corresponding files are named paperless_entity_mock_[correspondents | documenttypes | tags] respectively. These files must contain valid json content according to the Paperless-ngx API. This functionality is used to omit the need for a running paperless-ngx instance.

Usage

python -m paperlesslabelagent.agent

The agent will fetch your existing entities, process every PDF in INPUT_FOLDER, and then walk promts each proposal in the terminal, asking [y/n] questions as needed. Once every file is either confirmed or has exhausted its retry attempts, the run ends with a summary of what was and wasn't resolved.

For setting the execution strategy, for example, set STRATEGY=iterative in .env (or STRATEGY=iterative python -m paperlesslabelagent.agent) to use the iterative strategy.

Current limitations

  • Only PDF files are currently supported.
  • No functionality yet to push anything back to Paperless-ngx (no document upload or entity-creation) (open TODO).

License

Eclipse Public License 2.0

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages