Skip to content
View Kpraful's full-sized avatar
🎯
Focusing
🎯
Focusing

    Block or report Kpraful

    Block user

    Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

    You must be logged in to block users.

    Content in all repositories owned by your account will be closed.
    Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
    Report abuse

    Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

    Report abuse
    Kpraful/README.md

    Hey, I'm Praful 👋

    I build LLM systems that survive production.

    praful= {
    "role": "AI Software Engineer @ BCG X",
    "based_in": "Delhi NCR, India",
    "years": 4,
    "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
    "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
    "belief": "an agent that fails loudly beats an agent that guesses quietly",
    "debugging": "why it called the same tool six times",
    }

    PythonFastAPIAzureKubernetesPostgreSQLLangChain


    🧠 What I actually build

    Four things I have shipped to real tenants. Click any of them if you want the guts.

    🔌 Agentic AI, and a 39-tool MCP server

    Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

    • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
    • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
    • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
    • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
    🔍 Retrieval that holds up under questioning
    • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
    • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
    • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
    • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
    ⚙️ LLM-Ops, the unglamorous half
    • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
    • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
    • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
    • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
    🏗️ The platform underneath all of it
    • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
    • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
    • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
    • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

    📊 Receipts

    Numbers I can defend in an interview.

    What I builtWhat it moved
    🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
    ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
    📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
    🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
    🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
    🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
    🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

    🧰 Toolbox

    GenAI and LLM

    OpenAIAnthropicMCPLangChainRAGAgents

    Backend

    PythonFastAPIFlaskDjangoNode.js

    Cloud and DevOps

    AzureDockerKubernetesHelmTerraformGitHub Actions

    Data

    PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

    Frontend

    ReactTypeScriptViteRedux


    🏅 Badges that came with an exam

    • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
    • 🇮🇳 National Finalist, Smart India Hackathon 2020
    • 🧩 Top 27% on LeetCode, 300+ problems solved
    • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

    ☕ Off the clock

    • 🔬 Reading eval papers, then arguing with the benchmark
    • 🏗️ Rebuilding things that already work, but slower and with more logging
    • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
    • 🏆 Recovering hackathon person, still gets the itch every October

    📫 Find me

    LinkedInEmail

    Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

    Popular repositories Loading

    1. material-ui material-uiPublic

      Forked from mui/material-ui

      React components for faster and easier web development. Build your own design system, or start with Material Design.

      JavaScript

    2. git-test git-testPublic

      This is just for test

      Python

    3. it-cert-automation-practice it-cert-automation-practicePublic

      Forked from google/it-cert-automation-practice

      Google IT Automation with Python Professional Certificate - Practice files

      Python

    4. Cpp CppPublic

      C++

    5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

      Forked from FOSSEE/Scilab6-Test-Toolbox

      Scilab

    6. painlessMesh painlessMeshPublic

      Forked from gmag11/painlessMesh

      ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

      C++

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    Kpraful (Praful Katare) · GitHub
    Skip to content
    View Kpraful's full-sized avatar
    🎯
    Focusing
    🎯
    Focusing

      Block or report Kpraful

      Block user

      Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

      You must be logged in to block users.

      Content in all repositories owned by your account will be closed.
      Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
      Report abuse

      Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

      Report abuse
      Kpraful/README.md

      Hey, I'm Praful 👋

      I build LLM systems that survive production.

      praful= {
      "role": "AI Software Engineer @ BCG X",
      "based_in": "Delhi NCR, India",
      "years": 4,
      "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
      "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
      "belief": "an agent that fails loudly beats an agent that guesses quietly",
      "debugging": "why it called the same tool six times",
      }

      PythonFastAPIAzureKubernetesPostgreSQLLangChain


      🧠 What I actually build

      Four things I have shipped to real tenants. Click any of them if you want the guts.

      🔌 Agentic AI, and a 39-tool MCP server

      Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

      • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
      • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
      • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
      • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
      🔍 Retrieval that holds up under questioning
      • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
      • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
      • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
      • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
      ⚙️ LLM-Ops, the unglamorous half
      • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
      • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
      • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
      • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
      🏗️ The platform underneath all of it
      • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
      • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
      • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
      • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

      📊 Receipts

      Numbers I can defend in an interview.

      What I builtWhat it moved
      🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
      ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
      📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
      🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
      🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
      🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
      🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

      🧰 Toolbox

      GenAI and LLM

      OpenAIAnthropicMCPLangChainRAGAgents

      Backend

      PythonFastAPIFlaskDjangoNode.js

      Cloud and DevOps

      AzureDockerKubernetesHelmTerraformGitHub Actions

      Data

      PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

      Frontend

      ReactTypeScriptViteRedux


      🏅 Badges that came with an exam

      • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
      • 🇮🇳 National Finalist, Smart India Hackathon 2020
      • 🧩 Top 27% on LeetCode, 300+ problems solved
      • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

      ☕ Off the clock

      • 🔬 Reading eval papers, then arguing with the benchmark
      • 🏗️ Rebuilding things that already work, but slower and with more logging
      • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
      • 🏆 Recovering hackathon person, still gets the itch every October

      📫 Find me

      LinkedInEmail

      Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

      Popular repositories Loading

      1. material-ui material-uiPublic

        Forked from mui/material-ui

        React components for faster and easier web development. Build your own design system, or start with Material Design.

        JavaScript

      2. git-test git-testPublic

        This is just for test

        Python

      3. it-cert-automation-practice it-cert-automation-practicePublic

        Forked from google/it-cert-automation-practice

        Google IT Automation with Python Professional Certificate - Practice files

        Python

      4. Cpp CppPublic

        C++

      5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

        Forked from FOSSEE/Scilab6-Test-Toolbox

        Scilab

      6. painlessMesh painlessMeshPublic

        Forked from gmag11/painlessMesh

        ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

        C++

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Kpraful (Praful Katare) · GitHub
      Skip to content
      View Kpraful's full-sized avatar
      🎯
      Focusing
      🎯
      Focusing

        Block or report Kpraful

        Block user

        Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

        You must be logged in to block users.

        Content in all repositories owned by your account will be closed.
        Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
        Report abuse

        Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

        Report abuse
        Kpraful/README.md

        Hey, I'm Praful 👋

        I build LLM systems that survive production.

        praful= {
        "role": "AI Software Engineer @ BCG X",
        "based_in": "Delhi NCR, India",
        "years": 4,
        "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
        "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
        "belief": "an agent that fails loudly beats an agent that guesses quietly",
        "debugging": "why it called the same tool six times",
        }

        PythonFastAPIAzureKubernetesPostgreSQLLangChain


        🧠 What I actually build

        Four things I have shipped to real tenants. Click any of them if you want the guts.

        🔌 Agentic AI, and a 39-tool MCP server

        Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

        • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
        • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
        • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
        • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
        🔍 Retrieval that holds up under questioning
        • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
        • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
        • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
        • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
        ⚙️ LLM-Ops, the unglamorous half
        • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
        • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
        • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
        • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
        🏗️ The platform underneath all of it
        • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
        • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
        • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
        • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

        📊 Receipts

        Numbers I can defend in an interview.

        What I builtWhat it moved
        🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
        ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
        📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
        🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
        🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
        🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
        🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

        🧰 Toolbox

        GenAI and LLM

        OpenAIAnthropicMCPLangChainRAGAgents

        Backend

        PythonFastAPIFlaskDjangoNode.js

        Cloud and DevOps

        AzureDockerKubernetesHelmTerraformGitHub Actions

        Data

        PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

        Frontend

        ReactTypeScriptViteRedux


        🏅 Badges that came with an exam

        • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
        • 🇮🇳 National Finalist, Smart India Hackathon 2020
        • 🧩 Top 27% on LeetCode, 300+ problems solved
        • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

        ☕ Off the clock

        • 🔬 Reading eval papers, then arguing with the benchmark
        • 🏗️ Rebuilding things that already work, but slower and with more logging
        • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
        • 🏆 Recovering hackathon person, still gets the itch every October

        📫 Find me

        LinkedInEmail

        Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

        Popular repositories Loading

        1. material-ui material-uiPublic

          Forked from mui/material-ui

          React components for faster and easier web development. Build your own design system, or start with Material Design.

          JavaScript

        2. git-test git-testPublic

          This is just for test

          Python

        3. it-cert-automation-practice it-cert-automation-practicePublic

          Forked from google/it-cert-automation-practice

          Google IT Automation with Python Professional Certificate - Practice files

          Python

        4. Cpp CppPublic

          C++

        5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

          Forked from FOSSEE/Scilab6-Test-Toolbox

          Scilab

        6. painlessMesh painlessMeshPublic

          Forked from gmag11/painlessMesh

          ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

          C++

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Kpraful (Praful Katare) · GitHub
        Skip to content
        View Kpraful's full-sized avatar
        🎯
        Focusing
        🎯
        Focusing

          Block or report Kpraful

          Block user

          Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

          You must be logged in to block users.

          Content in all repositories owned by your account will be closed.
          Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
          Report abuse

          Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

          Report abuse
          Kpraful/README.md

          Hey, I'm Praful 👋

          I build LLM systems that survive production.

          praful= {
          "role": "AI Software Engineer @ BCG X",
          "based_in": "Delhi NCR, India",
          "years": 4,
          "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
          "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
          "belief": "an agent that fails loudly beats an agent that guesses quietly",
          "debugging": "why it called the same tool six times",
          }

          PythonFastAPIAzureKubernetesPostgreSQLLangChain


          🧠 What I actually build

          Four things I have shipped to real tenants. Click any of them if you want the guts.

          🔌 Agentic AI, and a 39-tool MCP server

          Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

          • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
          • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
          • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
          • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
          🔍 Retrieval that holds up under questioning
          • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
          • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
          • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
          • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
          ⚙️ LLM-Ops, the unglamorous half
          • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
          • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
          • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
          • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
          🏗️ The platform underneath all of it
          • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
          • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
          • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
          • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

          📊 Receipts

          Numbers I can defend in an interview.

          What I builtWhat it moved
          🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
          ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
          📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
          🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
          🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
          🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
          🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

          🧰 Toolbox

          GenAI and LLM

          OpenAIAnthropicMCPLangChainRAGAgents

          Backend

          PythonFastAPIFlaskDjangoNode.js

          Cloud and DevOps

          AzureDockerKubernetesHelmTerraformGitHub Actions

          Data

          PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

          Frontend

          ReactTypeScriptViteRedux


          🏅 Badges that came with an exam

          • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
          • 🇮🇳 National Finalist, Smart India Hackathon 2020
          • 🧩 Top 27% on LeetCode, 300+ problems solved
          • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

          ☕ Off the clock

          • 🔬 Reading eval papers, then arguing with the benchmark
          • 🏗️ Rebuilding things that already work, but slower and with more logging
          • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
          • 🏆 Recovering hackathon person, still gets the itch every October

          📫 Find me

          LinkedInEmail

          Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

          Popular repositories Loading

          1. material-ui material-uiPublic

            Forked from mui/material-ui

            React components for faster and easier web development. Build your own design system, or start with Material Design.

            JavaScript

          2. git-test git-testPublic

            This is just for test

            Python

          3. it-cert-automation-practice it-cert-automation-practicePublic

            Forked from google/it-cert-automation-practice

            Google IT Automation with Python Professional Certificate - Practice files

            Python

          4. Cpp CppPublic

            C++

          5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

            Forked from FOSSEE/Scilab6-Test-Toolbox

            Scilab

          6. painlessMesh painlessMeshPublic

            Forked from gmag11/painlessMesh

            ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

            C++

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Kpraful (Praful Katare) · GitHub
          Skip to content
          View Kpraful's full-sized avatar
          🎯
          Focusing
          🎯
          Focusing

            Block or report Kpraful

            Block user

            Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

            You must be logged in to block users.

            Content in all repositories owned by your account will be closed.
            Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
            Report abuse

            Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

            Report abuse
            Kpraful/README.md

            Hey, I'm Praful 👋

            I build LLM systems that survive production.

            praful= {
            "role": "AI Software Engineer @ BCG X",
            "based_in": "Delhi NCR, India",
            "years": 4,
            "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
            "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
            "belief": "an agent that fails loudly beats an agent that guesses quietly",
            "debugging": "why it called the same tool six times",
            }

            PythonFastAPIAzureKubernetesPostgreSQLLangChain


            🧠 What I actually build

            Four things I have shipped to real tenants. Click any of them if you want the guts.

            🔌 Agentic AI, and a 39-tool MCP server

            Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

            • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
            • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
            • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
            • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
            🔍 Retrieval that holds up under questioning
            • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
            • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
            • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
            • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
            ⚙️ LLM-Ops, the unglamorous half
            • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
            • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
            • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
            • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
            🏗️ The platform underneath all of it
            • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
            • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
            • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
            • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

            📊 Receipts

            Numbers I can defend in an interview.

            What I builtWhat it moved
            🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
            ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
            📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
            🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
            🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
            🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
            🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

            🧰 Toolbox

            GenAI and LLM

            OpenAIAnthropicMCPLangChainRAGAgents

            Backend

            PythonFastAPIFlaskDjangoNode.js

            Cloud and DevOps

            AzureDockerKubernetesHelmTerraformGitHub Actions

            Data

            PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

            Frontend

            ReactTypeScriptViteRedux


            🏅 Badges that came with an exam

            • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
            • 🇮🇳 National Finalist, Smart India Hackathon 2020
            • 🧩 Top 27% on LeetCode, 300+ problems solved
            • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

            ☕ Off the clock

            • 🔬 Reading eval papers, then arguing with the benchmark
            • 🏗️ Rebuilding things that already work, but slower and with more logging
            • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
            • 🏆 Recovering hackathon person, still gets the itch every October

            📫 Find me

            LinkedInEmail

            Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

            Popular repositories Loading

            1. material-ui material-uiPublic

              Forked from mui/material-ui

              React components for faster and easier web development. Build your own design system, or start with Material Design.

              JavaScript

            2. git-test git-testPublic

              This is just for test

              Python

            3. it-cert-automation-practice it-cert-automation-practicePublic

              Forked from google/it-cert-automation-practice

              Google IT Automation with Python Professional Certificate - Practice files

              Python

            4. Cpp CppPublic

              C++

            5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

              Forked from FOSSEE/Scilab6-Test-Toolbox

              Scilab

            6. painlessMesh painlessMeshPublic

              Forked from gmag11/painlessMesh

              ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

              C++

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Kpraful (Praful Katare) · GitHub
            Skip to content
            View Kpraful's full-sized avatar
            🎯
            Focusing
            🎯
            Focusing

              Block or report Kpraful

              Block user

              Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

              You must be logged in to block users.

              Content in all repositories owned by your account will be closed.
              Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
              Report abuse

              Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

              Report abuse
              Kpraful/README.md

              Hey, I'm Praful 👋

              I build LLM systems that survive production.

              praful= {
              "role": "AI Software Engineer @ BCG X",
              "based_in": "Delhi NCR, India",
              "years": 4,
              "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
              "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
              "belief": "an agent that fails loudly beats an agent that guesses quietly",
              "debugging": "why it called the same tool six times",
              }

              PythonFastAPIAzureKubernetesPostgreSQLLangChain


              🧠 What I actually build

              Four things I have shipped to real tenants. Click any of them if you want the guts.

              🔌 Agentic AI, and a 39-tool MCP server

              Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

              • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
              • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
              • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
              • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
              🔍 Retrieval that holds up under questioning
              • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
              • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
              • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
              • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
              ⚙️ LLM-Ops, the unglamorous half
              • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
              • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
              • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
              • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
              🏗️ The platform underneath all of it
              • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
              • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
              • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
              • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

              📊 Receipts

              Numbers I can defend in an interview.

              What I builtWhat it moved
              🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
              ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
              📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
              🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
              🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
              🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
              🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

              🧰 Toolbox

              GenAI and LLM

              OpenAIAnthropicMCPLangChainRAGAgents

              Backend

              PythonFastAPIFlaskDjangoNode.js

              Cloud and DevOps

              AzureDockerKubernetesHelmTerraformGitHub Actions

              Data

              PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

              Frontend

              ReactTypeScriptViteRedux


              🏅 Badges that came with an exam

              • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
              • 🇮🇳 National Finalist, Smart India Hackathon 2020
              • 🧩 Top 27% on LeetCode, 300+ problems solved
              • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

              ☕ Off the clock

              • 🔬 Reading eval papers, then arguing with the benchmark
              • 🏗️ Rebuilding things that already work, but slower and with more logging
              • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
              • 🏆 Recovering hackathon person, still gets the itch every October

              📫 Find me

              LinkedInEmail

              Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

              Popular repositories Loading

              1. material-ui material-uiPublic

                Forked from mui/material-ui

                React components for faster and easier web development. Build your own design system, or start with Material Design.

                JavaScript

              2. git-test git-testPublic

                This is just for test

                Python

              3. it-cert-automation-practice it-cert-automation-practicePublic

                Forked from google/it-cert-automation-practice

                Google IT Automation with Python Professional Certificate - Practice files

                Python

              4. Cpp CppPublic

                C++

              5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

                Forked from FOSSEE/Scilab6-Test-Toolbox

                Scilab

              6. painlessMesh painlessMeshPublic

                Forked from gmag11/painlessMesh

                ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

                C++

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Kpraful (Praful Katare) · GitHub
              Skip to content
              View Kpraful's full-sized avatar
              🎯
              Focusing
              🎯
              Focusing

                Block or report Kpraful

                Block user

                Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

                You must be logged in to block users.

                Content in all repositories owned by your account will be closed.
                Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
                Report abuse

                Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

                Report abuse
                Kpraful/README.md

                Hey, I'm Praful 👋

                I build LLM systems that survive production.

                praful= {
                "role": "AI Software Engineer @ BCG X",
                "based_in": "Delhi NCR, India",
                "years": 4,
                "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
                "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
                "belief": "an agent that fails loudly beats an agent that guesses quietly",
                "debugging": "why it called the same tool six times",
                }

                PythonFastAPIAzureKubernetesPostgreSQLLangChain


                🧠 What I actually build

                Four things I have shipped to real tenants. Click any of them if you want the guts.

                🔌 Agentic AI, and a 39-tool MCP server

                Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

                • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
                • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
                • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
                • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
                🔍 Retrieval that holds up under questioning
                • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
                • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
                • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
                • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
                ⚙️ LLM-Ops, the unglamorous half
                • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
                • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
                • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
                • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
                🏗️ The platform underneath all of it
                • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
                • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
                • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
                • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

                📊 Receipts

                Numbers I can defend in an interview.

                What I builtWhat it moved
                🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
                ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
                📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
                🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
                🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
                🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
                🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

                🧰 Toolbox

                GenAI and LLM

                OpenAIAnthropicMCPLangChainRAGAgents

                Backend

                PythonFastAPIFlaskDjangoNode.js

                Cloud and DevOps

                AzureDockerKubernetesHelmTerraformGitHub Actions

                Data

                PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

                Frontend

                ReactTypeScriptViteRedux


                🏅 Badges that came with an exam

                • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
                • 🇮🇳 National Finalist, Smart India Hackathon 2020
                • 🧩 Top 27% on LeetCode, 300+ problems solved
                • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

                ☕ Off the clock

                • 🔬 Reading eval papers, then arguing with the benchmark
                • 🏗️ Rebuilding things that already work, but slower and with more logging
                • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
                • 🏆 Recovering hackathon person, still gets the itch every October

                📫 Find me

                LinkedInEmail

                Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

                Popular repositories Loading

                1. material-ui material-uiPublic

                  Forked from mui/material-ui

                  React components for faster and easier web development. Build your own design system, or start with Material Design.

                  JavaScript

                2. git-test git-testPublic

                  This is just for test

                  Python

                3. it-cert-automation-practice it-cert-automation-practicePublic

                  Forked from google/it-cert-automation-practice

                  Google IT Automation with Python Professional Certificate - Practice files

                  Python

                4. Cpp CppPublic

                  C++

                5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

                  Forked from FOSSEE/Scilab6-Test-Toolbox

                  Scilab

                6. painlessMesh painlessMeshPublic

                  Forked from gmag11/painlessMesh

                  ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

                  C++

                , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Kpraful (Praful Katare) · GitHub
                Skip to content
                View Kpraful's full-sized avatar
                🎯
                Focusing
                🎯
                Focusing

                  Block or report Kpraful

                  Block user

                  Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

                  You must be logged in to block users.

                  Content in all repositories owned by your account will be closed.
                  Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
                  Report abuse

                  Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

                  Report abuse
                  Kpraful/README.md

                  Hey, I'm Praful 👋

                  I build LLM systems that survive production.

                  praful= {
                  "role": "AI Software Engineer @ BCG X",
                  "based_in": "Delhi NCR, India",
                  "years": 4,
                  "building": ["agent harnesses", "RAG that cites", "multi-tenant platforms"],
                  "stack": ["Python", "FastAPI", "Azure", "Postgres", "pgvector"],
                  "belief": "an agent that fails loudly beats an agent that guesses quietly",
                  "debugging": "why it called the same tool six times",
                  }

                  PythonFastAPIAzureKubernetesPostgreSQLLangChain


                  🧠 What I actually build

                  Four things I have shipped to real tenants. Click any of them if you want the guts.

                  🔌 Agentic AI, and a 39-tool MCP server

                  Built the natural-language layer of an enterprise platform: a 39-tool MCP server on FastMCP, exposing internal tools and data to LLM agents over Claude and Gemini.

                  • Tenant resolved from the authenticated Okta identity, never from caller input. 403 invariant on every cross-tenant request, with a security suite that proves it.
                  • Patched the tool decorator once so all 39 tools auto-wrap with error boundaries, correlation IDs and structured audit logs. Zero per-tool boilerplate.
                  • Guardrails that forbid the model from stating any number a tool did not return. No fine-tuning, no hallucinated market figures.
                  • Hand-rolled agent control loop: bounded 6-step tool calling, explicit stopping criterion, exceptions fed back as tool messages so the model recovers instead of crashing.
                  🔍 Retrieval that holds up under questioning
                  • Hybrid search fusing pgvector semantic retrieval with Postgres full-text via Reciprocal Rank Fusion, configurable 70/30 weighting plus document-type weighting.
                  • Multi-query expansion, token-by-token streaming, conversation memory, page-level citation extraction, all inside a 128K token budget.
                  • Document-type-aware recursive chunking with sentence-window and metadata enrichment into 1536-dimension embeddings.
                  • A vision path for presentations, then hierarchical DBSCAN clustering of embeddings into trend clusters with LLM-written summaries.
                  ⚙️ LLM-Ops, the unglamorous half
                  • Per-tenant multi-provider routing across OpenAI, Azure OpenAI and Gemini, with time-to-first-token fallback that re-routes mid-stream, a circuit breaker, and per-user and per-tenant inflight limits.
                  • 165+ versioned prompts resolved per-tenant-override then common-default, cached in Redis with hot reload, so subject-matter experts ship prompt changes without a redeploy.
                  • Every call traced in LangSmith and correlated into Datadog APM, with tiktoken cost accounting and per-provider context-window management.
                  • An LLM benchmarking harness scoring models on five quality dimensions plus latency, TTFT, throughput and cost, used to actually pick models.
                  🏗️ The platform underneath all of it
                  • Multi-tenant Azure self-service platform: app architecture, networking, Kubernetes, Container Apps Jobs. 10+ tenants, 10,000+ users.
                  • Per-tenant database isolation with encrypted connection strings, lazily created and dynamically budgeted pools, and ContextVar tenant propagation across async tasks. New tenants need no restart.
                  • Okta OIDC with spoof-proof group-to-tenant mapping, role-based endpoint guards, and JWT auto-refresh that never forces a re-login.
                  • Terraform: six reusable Azure modules standing up an isolated subscription per client, VNet, database and storage included, in about 30 minutes.

                  📊 Receipts

                  Numbers I can defend in an interview.

                  What I builtWhat it moved
                  🏢 Multi-tenant Azure platform10+ tenants, 10,000+ users, up to 5 clients per server
                  ⚡ Onboarding automation7-10 days ➜ 1-2 hours, and 6-month infra cost per client from $6,000 ➜ $1,200-$2,000
                  📄 Event-driven RAG pipeline50-100 docs per client, turnaround from weeks ➜ hours
                  🩺 LLM upload diagnostics agentCatches ~90% of errors pre-processing, debugging days ➜ minutes, answers in under 90s
                  🎯 GenAI synthetic survey panelRedesigned probabilistic aggregation, prediction error 47pt ➜ 10pt vs real respondents
                  🧮 Synthetic panel modeling coreSeeded k-means++ over 54-dim vectors, ~6,000 respondents ➜ ~250-300 prototypes, ~20x cheaper per question
                  🐘 Backend perf work (Infosys)Query optimization and pooling, latency down up to 30%

                  🧰 Toolbox

                  GenAI and LLM

                  OpenAIAnthropicMCPLangChainRAGAgents

                  Backend

                  PythonFastAPIFlaskDjangoNode.js

                  Cloud and DevOps

                  AzureDockerKubernetesHelmTerraformGitHub Actions

                  Data

                  PostgreSQLpgvectorMongoDBRedisSnowflakepandasNumPy

                  Frontend

                  ReactTypeScriptViteRedux


                  🏅 Badges that came with an exam

                  • 🎖️ Microsoft Certified: Azure Administrator Associate (AZ-104) and Azure Fundamentals (AZ-900)
                  • 🇮🇳 National Finalist, Smart India Hackathon 2020
                  • 🧩 Top 27% on LeetCode, 300+ problems solved
                  • 🎓 B.E. Computer Engineering, Bharati Vidyapeeth College of Engineering, Pune. CGPA 8.63

                  ☕ Off the clock

                  • 🔬 Reading eval papers, then arguing with the benchmark
                  • 🏗️ Rebuilding things that already work, but slower and with more logging
                  • 🧠 Convinced that most "the model is bad" bugs are actually retrieval bugs
                  • 🏆 Recovering hackathon person, still gets the itch every October

                  📫 Find me

                  LinkedInEmail

                  Always up for a conversation about agent infrastructure, retrieval, or why your LLM pipeline is slow. 🚀

                  Popular repositories Loading

                  1. material-ui material-uiPublic

                    Forked from mui/material-ui

                    React components for faster and easier web development. Build your own design system, or start with Material Design.

                    JavaScript

                  2. git-test git-testPublic

                    This is just for test

                    Python

                  3. it-cert-automation-practice it-cert-automation-practicePublic

                    Forked from google/it-cert-automation-practice

                    Google IT Automation with Python Professional Certificate - Practice files

                    Python

                  4. Cpp CppPublic

                    C++

                  5. Scilab6-Test-Toolbox Scilab6-Test-ToolboxPublic

                    Forked from FOSSEE/Scilab6-Test-Toolbox

                    Scilab

                  6. painlessMesh painlessMeshPublic

                    Forked from gmag11/painlessMesh

                    ESP8266 based mesh. This is a mirror copy of https://gitlab.com/painlessMesh/painlessMesh PLEASE ADD COMMENTS, ISSUES and PULL REQUESTS ON GITLAB so that all information is centralized.

                    C++