Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

ReportIQ Intelligence Engine

The Intelligence Engine is the AI backend for the ReportIQ platform. It is a Python-based microservice responsible for processing jobs, running local/cloud AI models, and managing vector embeddings.

Architecture

  • Language: Python 3.11
  • Framework: FastAPI (for API) + SQLAlchemy (for DB Access)
  • AI Framework: LangChain (for RAG & Embeddings)
  • Concurrency: Event Loop with SKIP LOCKED polling (No RabbitMQ required)

Project Structure

src/
├── domain/ # Pydantic contracts (data models)
│ └── models.py # IngestedContent, SearchResult, etc.
├── core/ # Shared, reusable modules
│ ├── ingestion/ # ONLY parsing logic
│ │ ├── parsers.py # File parsing (PDF, Excel, etc.)
│ │ └── utils.py # Helper utilities
│ ├── intelligence/ # ONLY AI/LLM logic
│ │ ├── embeddings.py # Ollama vector embeddings
│ │ └── llm_sql.py # Gemini SQL generation
│ ├── search/ # ONLY search logic
│ │ └── service.py # Hybrid semantic + keyword search
│ └── storage/ # ONLY MinIO logic
│ └── minio_service.py # File storage operations
├── features/ # Feature-specific modules (isolated)
│ ├── document_chat/ # PDF RAG feature
│ │ ├── service.py # Business logic for document chat
│ │ └── router.py # API endpoints (/documents/*)
│ └── data_analyst/ # Excel/SQL feature
│ ├── service.py # Business logic for SQL generation
│ └── router.py # API endpoints (/data/*)
├── db/ # Database layer
│ ├── models.py # SQLAlchemy ORM models
│ └── session.py # Database session management
├── config.py # Configuration (pydantic-settings)
├── main.py # FastAPI app entry point
└── worker.py # Background job processor
tests/
└── safety_net.py # Integration tests

Job Loop Logic

This service does not simply wait for HTTP requests. Instead, it actively polls the shared PostgreSQL database for work.

  1. Poll: Every 1 second, it checks job_queue for status='PENDING'.
  2. Lock: It uses SELECT ... FOR UPDATE SKIP LOCKED to atomically claim a job.
  3. Process: It routes the job to a specific handler based on job.type.
    • INGEST: Convert Text -> Vector (via Ollama/Nomic) -> Save to documents.
    • REPORT: Retrieve Vectors -> Summarize (via LLM) -> Save Report.
  4. Complete: It marks the job as DONE or FAILED.

Technical Stack

ComponentToolPurpose
DatabasePostgreSQL + pgvectorStores Jobs and Vector Embeddings
EmbeddingsOllama (nomic-embed-text)Converts text to 768-dim vectors locally
LLMOllama (mistral) / GeminiGenerates summaries and SQL queries
Configpydantic-settingsLoads .env variables securely
StorageMinIOObject storage for uploaded files

API Endpoints

Legacy Endpoints (Backward Compatible)

MethodEndpointDescription
GET/healthHealth check
POST/embedGenerate vector embedding
POST/parse_filesParse uploaded file
POST/searchHybrid document search
POST/chat/sqlNatural language to SQL

Feature Endpoints (New)

MethodEndpointDescription
POST/documents/parseParse uploaded document
POST/documents/embedGenerate embedding for text
POST/documents/searchSearch documents
POST/data/chat/sqlNatural language to SQL
GET/data/schemaGet available data schema keys

Quick Start

1. Environment Setup

Create a .env file in the root directory:

ENV_NAME=local
DATABASE_URL=postgresql://admin:adminpassword@localhost:5432/reportiq_db
OLLAMA_URL=http://localhost:11434
MINIIO_URL=localhost:9000
MINIIO_USERNAME=minioadmin
MINIIO_PASSWORD=minioadminpassword
MINIIO_SECURE=false
GEMINI_APIKEY=your_gemini_api_key

2. Install Dependencies

pip install -r requirements.txt

3. Run the Worker

This starts the background polling loop.

python -m src.worker

4. Run the API

Start the FastAPI server.

python -m src.main

Or with auto-reload for development:

uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

5. Run Safety Net Tests

Before starting any new feature, run the integration tests:

python tests/safety_net.py

Key Files

  • src/main.py: FastAPI application entry point (imports feature routers)
  • src/worker.py: Background job processor with polling loop
  • src/domain/models.py: Pydantic models for data contracts
  • src/features/*/service.py: Business logic for each feature
  • src/features/*/router.py: API endpoints for each feature
  • src/core/*/: Shared modules (ingestion, intelligence, search, storage)

Common Issues

  • 404 Model Not Found: The OLLAMA_URL is reachable, but the specific model (e.g., nomic-embed-text) hasn't been pulled on that server.
    • Fix: Run ollama pull nomic-embed-text on the server hosting Ollama.
  • Connection Refused: The Python script cannot see the Database.
    • Fix: If running locally, ensure DATABASE_URL uses localhost. If running in Docker, use db.

About ReportIQ

ReportIQ is an enterprise AI platform designed to transform raw company data into actionable intelligence. Our mission is to enable organizations to plug in diverse data sources—including databases, emails, spreadsheets, and documents—and instantly generate reports, summaries, and answers without manual effort.

Core Capabilities

  • Unified Data Ingestion: Seamlessly ingest and normalize data from disparate sources such as SQL databases, PDF documents, Excel spreadsheets, and email servers.
  • Hybrid AI Intelligence: Intelligently route queries between efficient local models for privacy and powerful cloud models for complex reasoning.
  • Automated Reporting: Generate scheduled executive summaries, operational dashboards, and anomaly reports automatically.
  • Enterprise Security: Built with strict multi-tenancy, row-level security, and identity propagation to ensure data isolation and compliance.

Technology Stack Overview

  • Core Platform: Java 21 / Spring Boot (Orchestration, Security, API Gateway)
  • Intelligence Engine: Python / FastAPI / LangChain (RAG, Vector Processing, LLM Inference)
  • Frontend Console: TypeScript / SvelteKit (Reactive User Interface)
  • Data Layer: PostgreSQL 17 (Relational Data) + pgvector (Semantic Search) + pgai (In-Database AI)
  • Infrastructure: Docker Compose (Local Development) / Docker Swarm (Deployment)

Key Repositories

Core Services

  • reportiq-core: The central backend service handling authentication, API gateway functions, and business logic.
  • reportiq-intelligence: The AI worker service responsible for document parsing, vector embedding, and LLM interactions.
  • reportiq-console: The web-based administration and user interface.

Shared & Infrastructure

  • reportiq-infra: Infrastructure as Code (IaC), Docker configurations, and deployment scripts.

Contact & Collaboration

This organization is currently in active development. For inquiries regarding access, security, or contribution guidelines, please contact the engineering team directly.

About

The AI worker service responsible for document parsing, vector embedding, and LLM interactions

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages