Skip to content
View omarmohammed271's full-sized avatar

Block or report omarmohammed271

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
omarmohammed271/README.md

👋 Hi, I'm Omar Mohammed

Senior Data Engineer | Real-Time & Batch Data Architect


🧱 About Me

  • Passionate about building scalable data pipelines using Spark, Kafka, and modern data platforms.
  • Experienced in both cloud-native architectures (AWS, Azure) and on-premise systems.
  • Strong believer in the power of clean code, observability, and data quality.

💼 What I Do

  • Design and develop ETL and ELT pipelines (batch + streaming).
  • Implement data lakehouse architectures (Bronze / Silver / Gold stages).
  • Build real-time processing systems using Spark Structured Streaming & Kafka.
  • Automate workflows and scheduling with Airflow.
  • Optimize analytics databases (e.g. ClickHouse) for fast query performance.
  • Create CI/CD pipelines for data solutions (using GitHub Actions or similar).

🌐 Tech Stack

DomainTechnologies
Data ProcessingPySpark, Spark SQL, Delta Lake
StreamingApache Kafka, Spark Structured Streaming
Workflow OrchestrationApache Airflow
Data StorageS3 / ADLS, Delta / Parquet
AnalyticsClickHouse, PostgreSQL
CloudAWS, Azure
CI / CDGitHub Actions
LanguagesPython, SQL

🚀 Featured Projects

Here are some of my key repositories (feel free to click and explore):

  • [Real-Time Processing Pipeline] — A Kafka → Spark Streaming system with schema validation and data quality checks.
  • [Lakehouse Architecture Demo] — Multi-layer (Bronze / Silver / Gold) data lakehouse built with Delta Lake.
  • [Airflow Data Workflows] — End-to-end DAGs for ingestion, transformation, and orchestration.
  • [Analytics in ClickHouse] — Setup for real-time analytics using ClickHouse materialized views.
  • [CI/CD for Data Jobs] — GitHub Actions to test, build, and deploy data workloads.

📈 GitHub Stats

Omar’s GitHub stats


📫 Get in Touch


⚡ Fun Facts

  • I love optimizing pipelines — every millisecond matters.
  • Outside work: I enjoy reading about distributed systems and data infrastructure.
  • Lifelong learner: currently exploring feature stores and ML data platforms.

🌐 Socials:

LinkedIn

💻 Tech Stack:

🚀 Tech Stack

🧱 Data Engineering

SparkKafkaAirflowClickHouseTrinoHadoop


🐍 Programming

PythonScalaJavaSQL


☁ Cloud & DevOps

AWSAzureDockerKubernetesGitHub ActionsGitGitHubGitLab


🗄 Databases & Warehousing

PostgreSQLMySQLMongoDBRedisDelta LakeS3


📊 Data Science / ML Tools

NumPyPandasscikit-learnPyTorchMatplotlib

📊 GitHub Stats:



🏆 GitHub Trophies

🔝 Top Contributed Repo


Pinned Loading

  1. Advanced-E-commercial-DjangoAdvanced-E-commercial-DjangoPublic

    Advanced E-commercial website

    HTML