A Streamlit-based app with a FastAPI backend for extracting structured data (text, images, tables) from websites and PDFs. Processed data is stored in AWS S3 and rendered in a markdown-standardized format. APIs are deployed on Google Cloud Run Service
dockeraws-s3pdf-converterpython3scrapydiffbotwebscrapingweb-data-extractiondiffbot-apibeautifulsoup4pymupdfpdf-document-processorgoogle-cloud-runstreamlitazure-document-intelligencedoclin
-
Updated
Jan 31, 2025 - Jupyter Notebook