View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
View irfanalidv's full-sized avatar

Organizations

@brainsfeed@re-sources-io

Block or report irfanalidv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
irfanalidv/README.md
Irfan Ali — Senior AI Engineer. I make LLM systems reliable in production.

WebsiteLinkedInPyPIORCIDEmail


Senior AI engineer, seven-plus years. I work on the part that's hard after the demo works: retrieval that returns the right answer, agents that degrade loudly instead of silently, evaluation harnesses that catch regressions, and cost you can actually account for.

Founding AI engineer at Kuration AI. First AI hire at Schneider Electric's Luminous R&D, reporting to the CTO. First data hire at brainsfeed, where I led a distributed team of ten.

Open to senior AI IC roles — Bengaluru or remote.irfan.ali@datacortex.in

Experience

RoleOrgWhen
Independent AI Engineer (self-employed)DataCortex IQ · IndiaDec 2025 – present
Founding AI EngineerKuration AI · Hong Kong2024–2025
Senior Manager – Data & AI, R&D · first AI hire, reported to CTOLuminous Power Technologies (Schneider Electric) · India2023–2024
Data Analytics & Automation AssociateLynk · India2022–2023
Head of Data & Analytics · first data hire, led team of 10+brainsfeed · Hong Kong2018–2022

Since December 2025 I've worked independently, self-funding a focused build phase on production LLM infrastructure — evaluation, reliability, retrieval, and multi-provider routing — and open-sourcing most of it. Pre-revenue by design and by outcome. The systems below came out of it.

Selected engineering work

Full writeups at datacortex.in. Every number below has a public reproduce path.

Company-intelligence extraction pipeline — Five stages: search → crawl → extract → deduplicate → verdict. FastAPI, DSPy, Pydantic strict JSON Schema, confidence-scored verdicts with human-review flagging, per-call cost metering. Pass 3 runs sequentially so each verdict sees already-confirmed products.

Frozen eval, 2026-09-01, 20/20 domains scored: macro P / R / F1 = 0.771 / 0.869 / 0.795. The writeup is mostly about the misses — including an optimizer run that made the metric worse and got reverted, and a search provider that returned HTTP 400 on empty credits so the pipeline scored zeros instead of failing loudly.

Case study · Eval writeup · Labels + rescore script

FastAPIDSPyPydanticLangSmith

Voice check-in system — Idempotent webhook ingest, deterministic safety detection on the raw transcript before any LLM call, then a Groq → Hugging Face → heuristic fallback chain. Cross-session context comes from pgvector retrieval over prior calls, not a longer prompt.

Case study

Next.jsBolnaGroqpgvector

FMCG trade-operations app — Party ledgers with VAT/PAN, billing, collections, credit limits, godown stock, field-visit logging, phone-width UI. Built for distributors in Nepal. No LLM in the critical path — correctness here is referential integrity and role checks, and the writeup is honest about where app-level roles should have been database policies.

Case study

Next.js 15TypeScriptSupabase

Open source

Twelve published Python libraries. The only numbered result is linked to its benchmark file.

LibraryWhat it does
RAGNav · srcHybrid BM25 + dense retrieval with RRF fusion. Navigation-first RAG for long documents. R@3 0.956 over 500 SQuAD questions — script and results in-repo
ragfallback · srcStop RAG from failing silently. Query rewriting, retrieval confidence scoring, fallback strategies, retry logic
AgentEnsemble · srcMulti-agent orchestration. ReAct, Swarm, Pipeline, Debate, WorkflowGraph. Routing, planning, RAG, cost tracking
nepal-gov-agent · srcAgentic RAG on Nepal government policy and legal documents. Hybrid retrieval with citations. Nepali + English
AgentCare · srcVoice AI for healthcare. Call intake, structured extraction, missing-data recovery, appointment orchestration
scrapeflow-py · srcPlaywright scraping. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection
AskPandas · srcNatural-language queries on CSV via local LLMs. No API keys, no cloud
lingo-nlp-toolkit · srcLightweight NLP utilities bridging classic pipelines and transformer-ready workflows
PyroChain · srcAgentic feature engineering. PyTorch + LangChain agents for multimodal extraction
toxic-comment-classifier · srcDeep-learning toxicity detection with per-category scores
socialmediaextractor · srcExtract social media profile links from websites
trustpilot-scraper · srcScrape Trustpilot reviews into structured output

All packages: pypi.org/user/irfanalidv

Applied research

  • Cross-validation framework for mental-health AI on MentalChat16K — BERT and neural networks · IJAINN, Dec 2025 · DOI
  • Neural-symbolic topic evolution on Yelp reviews — multi-aspect temporal topic modelling · IJAINN, Oct 2025 · DOI

ORCID: 0000-0003-0022-3047

Contact

irfan.ali@datacortex.in · LinkedIn · datacortex.in

Pinned Loading

  1. ragfallbackragfallbackPublic

    ragfallback is a Python library that prevents silent RAG failures — chunk quality, retrieval fallback, adaptive querying, and answer evaluation in one package.

    Python 2

  2. AgentEnsembleAgentEnsemblePublic

    AgentEnsemble is a Production-ready multi-agent orchestration for Python. ReAct, Swarm, Pipeline, Debate, Router, Planner, WorkflowGraph. Observability, cost tracking, human-in-loop. Structured out…

    Python 1

  3. AskPandasAskPandasPublic

    AI-powered data engineering and analytics assistant for querying CSV data using natural language—locally and intelligently

    Python 1

  4. scrapeflow-pyscrapeflow-pyPublic

    Production-ready web scraping engine on Playwright. LLM extraction, hybrid selectors, session persistence, rate limiting, anti-detection.

    Python 1

  5. AgentCareAgentCarePublic

    Python framework for voice-AI workflows: healthcare front-desk booking, care coordination, follow-up, and workplace burnout check-ins.

    Python 1

  6. RAGNavRAGNavPublic

    RAGNav is a Hybrid RAG retrieval — BM25 + embeddings + document graph. Runs fully offline. SQuAD R@3: 0.956, zero API calls.

    Python 1