Skip to content
View jmdu99's full-sized avatar

Block or report jmdu99

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jmdu99/README.md

Hi there 👋, I'm Jose

Data Engineer — Turning messy data into clarity for projects with real impact.
I work with purpose-driven teams to build data systems they can trust.

🛠 What I Do

  • 📊 Centralise scattered data into a single source of truth
  • ⚙️ Automate cleaning & validation for always-ready data
  • 🚀 Design efficient ETL/ELT pipelines (Airflow, dbt, Spark…)
  • 📈 Build solid foundations for BI, ML & GenAI
  • ⏱ Create real-time dataflows when speed matters

💻 Tech Stack

Core Skills & Tooling

PythonSQLBashGitGitHubPoetryPylintPandasNumPy

Ingestion, Orchestration & Processing

Apache AirflowCloud Composer (GCP)MWAA (AWS)dbtFivetranAirbytePrefectApache SparkPySparkApache BeamDataflow (GCP)Dataproc (GCP)Spark Structured StreamingApache KafkaGoogle Pub/SubApache NiFiWeb scraping

Data Platforms & Storage

Amazon S3Google Cloud StorageParquetBigQuerySnowflakeAmazon RedshiftAmazon AthenaPostgreSQLMongoDBCassandraClickHouse

Cloud & DevOps

Amazon EC2Google Compute EngineTerraform (IaC)DockerDocker ComposeGitHub Actions (CI/CD)IAM / RBAC

ML, NLP & Knowledge Graphs

Generative AILarge Language ModelsOpenAI APILangChain (RAG)Hugging Face TransformersNLTKspaCyscikit-learnPyTorchTensorFlowSPARQLAWS SageMaker

Analytics & Visualization

MatplotlibSeabornPlotlyAmazon QuickSightApache Superset

🏆 GitHub Trophies

trophy

Pinned Loading

  1. Hybrid-Fitness-Data-Pipeline-Batch-StreamingHybrid-Fitness-Data-Pipeline-Batch-StreamingPublic

    This project demonstrates a full hybrid fitness data pipeline combining real-time streaming (Kafka + MongoDB) with scheduled batch enrichment and loading (Prefect + Redshift). Dashboards are built …

    Python

  2. Hybrid-Nutrition-Data-Pipeline-Batch-StreamingHybrid-Nutrition-Data-Pipeline-Batch-StreamingPublic

    This project simulates a real-time and batch data pipeline for food item enrichment and nutritional analytics. It demonstrates a modern architecture that uses Kafka for streaming ingestion, Cassand…

    Python

  3. dbpedia/DBpedia-Spotlight-Dashboarddbpedia/DBpedia-Spotlight-DashboardPublic

    An integrated statistical information tool from the Wikipedia dumps and the DBpedia Extraction Framework artifacts

    Python 1 1

  4. Data-Processes-assignmentData-Processes-assignmentPublic

    COVID-19 survival analysis of a dataset and prediction using Python (sklearn, pandas, numpy, matplotlib, lifelines, mlxtend, joblib)

    Python 1

  5. Spark-Practical-WorkSpark-Practical-WorkPublic

    Big Data: Spark Practical Work First Semester 2021/2022

    Scala

  6. Graph-Analysis-Social-NetworksGraph-Analysis-Social-NetworksPublic

    Assignments made during the Graph Analysis and Social Networks course using Tweepy, NetworkX and NLTK.

    Jupyter Notebook 1