Skip to content
View arijitroy003's full-sized avatar

Organizations

@redhat-data-and-ai

Block or report arijitroy003

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
arijitroy003/README.md

Arijit Kumar Roy — Data and AI systems built for production

Portfolio · LinkedIn · Email

I turn expensive, ambiguous data operations into dependable production systems: agents that shorten incident response, platforms that make governed data self-service, and automation that gives senior engineers their time back.

Today I am a Senior Software Engineer in Data & AI Platform Engineering at Red Hat. Across eight years, I have built for regulated enterprise platforms and consumer products serving up to 120M+ users and 500M events per day.

What I build

01 / Agentic AI

Production MCP and LangChain systems for investigation, release orchestration, metadata intelligence, and operational decision support—with observability and human handoffs designed in.

02 / Data platforms

Self-service data products on Snowflake, Databricks, dbt, Spark, and Delta Lake, with governance, quality, lineage, and cost controls embedded in the paved road.

03 / Platform engineering

Cloud-native control planes and developer workflows built with Python, Go, Kubernetes, OpenShift, GitOps, Terraform, and pragmatic automation.

Systems I have shipped

SystemWhat changedScale and stack
Data Reliability AgentAutomated first-line pipeline triage and reduced initial investigation from ~38 minutes to ~4 minutes.MCP, LangChain, OpenShift, Langfuse
Release AssistantOrchestrated governed releases and saved 1,000+ lead-engineer hours annually.Python, GitOps, Snowflake, 150+ compliant data products
Self-service Data MeshReplaced legacy Redshift/Starburst paths and reduced infrastructure cost by $200K+/year.OpenShift, dbt, Snowflake, Kubernetes
Consumer AI at scaleBuilt GenAI search, recommendations, and conversational systems for Tata Neu and Beem.120M+ users, 500M events/day, 12 Indic languages

Open source, upstream

I contribute correctness fixes, security hardening, CI improvements, typing support, tests, and documentation across the data and AI ecosystem.

60 merged upstream pull requests across 28 repositories, including the DuckDB, vLLM, dbt Labs, Red Hat, llm-d, Apache, and LangChain communities.

Latest merged upstream pull requests

Selected engineering contributions

Current lab

Small, public experiments where I explore agent interfaces, developer tooling, data workflows, and useful automation.

ProjectWhat it exploresLanguage
linkedin-mcp-serverLinkedIn automation MCP server wrapping the unofficial linkedin-apiPython
snap-a-miroConvert whiteboard photos into interactive Miro boards using AI vision analysisJavaScript
datadiffHigh-performance CLI tool for semantic diffing of tabular data (CSV, Excel, Parquet, JSON) with Git integrationRust
flight-trackerLocal flight price tracker with web UI - supports Amadeus & Skyscanner APIs, daily price monitoring, Indian market optimizedPython
Recent public activity

Enterprise contributions

Most of my production work ships to Red Hat's private GitLab. This activity graph provides the missing context that a public GitHub contribution graph cannot.

Red Hat GitLab contribution activity from July 2025 through July 2026

Technical toolkit

languages Python · Go · Rust · SQL · TypeScript
data Snowflake · dbt · Databricks · Spark · Delta Lake · Kafka · Airflow
ai systems LangChain · MCP · OpenAI · Claude · Mistral · Vector DBs · Langfuse
platform Kubernetes · OpenShift · GitOps · Terraform · Docker · AWS · Azure · GCP

MCA, Jadavpur University · Distributed systems and information retrieval research at ISI Kolkata


Building a serious data platform or production AI system?
Start a conversation · Explore my work

Project, activity, and upstream contribution data refresh automatically through GitHub Actions.

Pinned Loading

  1. arijitroy003.github.ioarijitroy003.github.ioPublic

    Personal portfolio website showcasing my work in Data Engineering, AI/ML, and Software Development

    JavaScript

  2. datadiffdatadiffPublic

    High-performance CLI tool for semantic diffing of tabular data (CSV, Excel, Parquet, JSON) with Git integration

    Rust

  3. linkedin-mcp-serverlinkedin-mcp-serverPublic

    LinkedIn automation MCP server wrapping the unofficial linkedin-api

    Python

  4. snap-a-mirosnap-a-miroPublic

    Convert whiteboard photos into interactive Miro boards using AI vision analysis

    JavaScript

  5. duckdbduckdbPublic

    Forked from duckdb/duckdb

    DuckDB is an analytical in-process SQL database management system

    C++

  6. genai-toolboxgenai-toolboxPublic

    Forked from googleapis/mcp-toolbox

    MCP Toolbox for Databases is an open source MCP server for databases.

    Go