Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - 92kareeem/Sentinel-Agentic-RAG: A self-healing, guardrailed agentic RAG platform on AWS · GitHub
Skip to content

Repository files navigation

Sentinel — Document Self-Healing Agentic RAG

A guardrailed, self-healing retrieval-augmented generation platform. Built end-to-end and deployed on AWS free tier.

Stack: FastAPI · LangGraph · Hybrid FAISS + BM25 (RRF) · Groq (openai/gpt-oss-20b routing/critic, 120b escalation) · AWS Lambda + DynamoDB + S3 + CloudFront · pytest · GitHub Actions


What it does

Most RAG systems are one-shot: retrieve, generate, ship. When retrieval misses or the model hallucinates, the user sees the failure.

Sentinel wraps the pipeline in a LangGraph agent that grades its own output and repairs it before responding.

 ┌─────────┐
query ──────▶│ router │─── simple? ──▶ direct answer
└────┬────┘
│ complex
▼
┌─────────┐ ┌──────────────┐
│retriever│───▶│ synthesiser │
└─────────┘ └──────┬───────┘
▼
┌─────────┐
│ critic │
└────┬────┘
│ low-confidence
▼
┌─────────┐
│ repair │──── loop back to retriever
└─────────┘
  • Router — Groq 8B classifies query complexity and routes cheaply.
  • Hybrid retriever — FAISS (dense, IndexFlatIP) + BM25 (sparse) fused with Reciprocal Rank Fusion (k=60). Sparse recall for exact terms, dense recall for meaning.
  • Synthesiser — grounded strictly on retrieved chunks; document text is fenced and labelled untrusted so content inside an uploaded file cannot act as an instruction.
  • Critic — evaluates the answer against retrieved context. Faithfulness and relevance scored.
  • Repair loop — on low confidence, rewrites the query and re-retrieves. Bounded to prevent runaway loops.
  • Guardrails — input and output. Blocks prompt-injection patterns, PII leakage, off-topic drift.
  • Full request tracing — every node emits structured logs; traces stored in DynamoDB for replay and debugging.

The point isn't the framework choices. The point is the platform grades itself, catches its own failures, and only ships answers it can defend.


Status

  • Ingestion pipeline and hybrid index (make ingest)
  • LangGraph agent: router → retriever → synthesiser → critic → grounding → repair → refusal
  • FastAPI service, guardrail chain, 90-test suite (unit + HTTP contract + end-to-end)
  • Document registry with explicit lifecycle, ownership and tenant isolation
  • Page-aware PDF pipeline with typed failure modes (encrypted / scanned / corrupt)
  • Atomic, versioned index publication safe under concurrent uploads
  • AWS deploy: Lambda container, DynamoDB, S3, CloudFront (infra/deploy.sh)
  • Evaluation harness (faithfulness / retrieval-hit / refusal-rate), GitHub Actions CI
  • TypeScript frontend

Known limitations

These are deliberate scope boundaries, not oversights:

  • No OCR. Image-only/scanned PDFs are detected and rejected with UNSUPPORTED_SCANNED_DOCUMENT rather than indexed as empty.
  • Answers are not token-streamed./v1/query returns one JSON body after the graph completes; the UI renders it progressively (labelled as such, not as streaming).
  • Ingestion is synchronous. Fine for the 1 MB / 200-page limit this targets. A durable queue (S3 event → SQS → worker) is the right shape beyond that; the previous in-process daemon thread was removed because Lambda freezes on return.
  • Index publication is single-writer. Concurrent uploads are serialized and retried, which is correct but not high-throughput.
  • Cross-process index locking is advisory (compare-and-set on the version pointer). Correct for one Lambda instance; multi-instance concurrent writes need the DynamoDB-conditional lock described in infra/aws_setup.md.

Architecture decisions

DecisionChoiceWhy
OrchestrationLangGraphExplicit state machine; conditional edges make the repair loop trivial
Vector storeFAISS IndexFlatIP, immutable versioned artifacts on S3Exact search, zero servers; versioning makes publication atomic
Document identityServer-generated uuid + DynamoDB registryFilenames are neither unique across users nor stable across re-uploads
Sparse retrievalBM25Exact-term recall the dense index misses
FusionReciprocal Rank Fusion (k=60)No score calibration needed across dense/sparse
LLM providerGroqFast inference, generous free tier
Small/large splitgpt-oss-20b (router, critic) + 120b (escalation)Most cost lives on the small model; escalate only when repair needs it
ServingAWS Lambda container image behind API GatewayCold start acceptable for demo; scales to zero; free tier
StateDynamoDB (traces, API keys)Serverless, single-digit-ms reads, no schema migrations
AuthAPI Gateway usage plans + hashed keys in DynamoDBTwo layers of protection, no Cognito overhead
FrontendReact + TypeScript on CloudFrontStatic hosting, cheap, edge-cached
Testspytest, GitHub Actions on pushIngestion, retrieval, agent nodes, guardrails, end-to-end

Repo layout

backend/ FastAPI service, LangGraph nodes, guardrails, retrieval
frontend/ React + TypeScript chat UI
infra/ IaC: Lambda, API Gateway, DynamoDB, S3, CloudFront
evals/ Evaluation harness, golden dataset, report generator
docker/ Lambda container image
docs/ Sample documents for the demo corpus
.github/workflows/ CI — lint, test, eval on push

Quickstart (local)

# Environment: venv must live outside cloud-synced folders (OneDrive corrupts native DLLs)# and be built from a standalone CPython (conda-derived venvs break torch DLL init).
uv venv C:/venvs/sentinel --python 3.12
uv pip install -e ".[dev]" --python C:/venvs/sentinel/Scripts/python.exe
cp .env.example .env # fill GROQ_API_KEY
make ingest # build index/ from ./docs, runs smoke test
make test# full pytest suite
make eval# run evaluation harness → evals/report.md
make serve # local FastAPI on :8000

Deploy

export AWS_ACCOUNT_ID=... GROQ_API_KEY=...
make deploy # == bash infra/deploy.sh

infra/deploy.sh is idempotent and creates/updates: S3 buckets (+ lifecycle, CORS), the four DynamoDB tables, the least-privilege IAM role, the ECR repo and image, the Lambda function and its URL, and finally syncs the local index — publishing the version directory before the pointer so a cold-starting Lambda never reads a torn index. It ends with a /healthz smoke test.

The container image fetches the quantized MiniLM at build time rather than copying a gitignored models/onnx/, so a fresh clone builds (this is enforced by a CI job).

Auth is a hashed API key in DynamoDB, checked per request. The frontend ships only a non-privileged, quota-limited demo key; the admin key must never be built into it.


Why this project

Every AI Engineer job ad in 2026 mentions RAG, agents, evals and guardrails as bullet points on a wishlist. This is what those bullet points actually look like when they meet each other in production: a system that routes cheaply, retrieves hybrid, grades itself, repairs when it fails, and refuses to answer when it can't be sure. Every architecture decision above is a real trade-off I made.

Built by Syed Abdul Kareem Ahmed.

About

A self-healing, guardrailed agentic RAG platform on AWS

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages