Searchable Markdown and source PDF of Malaysia’s Royal Commission of Inquiry report on Lembaga Tabung Haji, with page-linked sections and findings guide.
-
Updated
Jul 30, 2026 - Python
Searchable Markdown and source PDF of Malaysia’s Royal Commission of Inquiry report on Lembaga Tabung Haji, with page-linked sections and findings guide.
AI-powered document scanner that automatically detects, corrects perspective, and enhances scanned documents from photos using OpenCV
Digitizing 570K Oklahoma oil & gas well records (1911-2024): OCR + CV + U-Net extraction to a live, human-verifiable map
图片转文字 · Chinese OCR Skill for Claude Code — TIFF/PNG/JPG to clean Markdown with watermark removal, page sorting, auto-correction & self-improving pipeline
Python pipeline for OCR+LLM document digitization
AI-powered full-stack web application for digitizing handwritten and printed documents using PaddleOCR, CRNN, OpenCV, and Ollama (Qwen2.5) for intelligent OCR correction.
OCR System for Extracting Text from Scanned PDF Documents using PaddleOCR and Streamlit
Archived LCT 2025 prototype: FastAPI pipeline for segmenting and recognizing Russian historical document scans with TrOCR/PaddleOCR.
Interactive OCR system with real-time correction using Tesseract and confidence-based filtering.
Add a description, image, and links to the document-digitization topic page so that developers can more easily learn about it.
To associate your repository with the document-digitization topic, visit your repo's landing page and select "manage topics."