📊 Data Scientist • 🤖 Machine Learning • 🚀 Production ML
🇲🇽 Based in CDMX | LinkedIn | Portfolio
Hi! I'm Fernando — I like turning raw data into cool stuff. I build data-powered solutions that help businesses make smarter decisions. Currently geeking out on RAG 🔍 and AWS ☁️ to build smarter, scalable AI systems.
🚀 Latest production model: Real estate valuation API (LightGBM, R² = 0.904) deployed via Django + Docker. Cut valuation time from 4 hours → 39 seconds (-99.73%).
I'm all about:
- Models that ship, not just notebooks
- Clean pipelines and reproducible code
- Real business impact over fancy metrics
Languages & Core
Data & ML
Deep Learning
Production
- 📉 Gradient Descent optimization
- 🌲 Gradient Boosting: XGBoost, LightGBM, CatBoost
- 🔄 Feature encoding (One-Hot, Label, Target)
- ⚖️ Feature scaling (Standard, MinMax, Robust)
- 🧪 EDA, feature engineering, modeling
- 🎯 Cross-validation & hyperparameter tuning
- 🧬 ML pipelines & API deployment
- 🖼️ Computer Vision: CNNs, ResNet50, Transfer Learning
- 📝 NLP: BERT embeddings, TF-IDF, spaCy
- 🔤 Text classification: Logistic Regression, Naive Bayes
- 🔁 Make, n8n, Webhooks, REST APIs
End-to-end ML system for property valuation in the Mexican market. Deployed as a REST API with automated weekly retraining.
Stack: LightGBM · scikit-learn · Django · Docker · VPS Impact: R² = 0.904 · 4h → 39s (-99.73%) · 80% confidence intervals
Churn prediction system for a telecom operator. Built 9 models, tuned with GridSearchCV/RandomizedSearchCV, and selected an ensemble for stability.
Stack: XGBoost · CatBoost · RandomForest · scikit-learn Result: AUC-ROC 0.847 (ensemble Top 3) · 5-fold stratified CV Business impact: Identified ~$2.7M USD in churn losses · prioritized retention list with 5,174 active customers in 3 risk segments
