View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
View SamHavocH's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SamHavocH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamHavocH/README.md

Linux-firstData EngineeringTerminal Workflows

Sam Havoc

Linux-first Data Engineer focused on distributed data systems, automation, and infrastructure-minded software.

I build practical systems around PySpark, Kafka, Airflow, Docker, and Python: pipelines that can be observed, replayed, debugged, and operated from a terminal without ceremony. My work leans toward production data platforms, developer tooling, TUI interfaces, and clean operational habits.

data pipelines / linux systems / terminal tooling / distributed workflows

Featured Work

ProjectFocusEngineering angle
pyspark-medallion-platformPySpark, lakehouse architecture, medallion layersBatch data platform with bronze/silver/gold modeling, validation, and reproducible local execution.
airflow-dbt-warehouseAirflow, dbt, warehouse orchestrationScheduled ELT workflows with explicit dependencies, transformation boundaries, and maintainable DAG design.
streaming-kafka-data-qualityKafka, streaming, data qualityEvent-driven validation pipeline for catching malformed records before they poison downstream systems.
tui-news-readerPython, TUI, terminal UXKeyboard-driven terminal application built for fast reading, filtering, and low-friction daily use.
document-quality-pipelineAutomation, parsing, quality checksDocument ingestion and validation workflow focused on repeatability, traceability, and clean outputs.
.dotfilesLinux, shell, editor workflowPersonal operating environment for terminal-first development, automation, and fast context switching.

Technical Stack

Languages
Python, SQL, Bash

Data Engineering
PySpark, Apache Spark, Kafka, Airflow, dbt, ETL/ELT, data quality, medallion architecture

Infrastructure
Linux, Docker, Git, GitHub Actions, local-first development environments, reproducible workflows

Linux Tooling
tmux, Neovim, zsh, shell scripting, dotfiles, CLI automation, TUI applications

Observability
Structured logs, pipeline checks, failure visibility, operational metrics, debuggable data flows


Current Focus

  • Building data platforms that are easy to run locally and realistic enough to resemble production.
  • Deepening distributed systems fundamentals around streaming, partitioning, retries, and backpressure.
  • Improving pipeline observability with better validation, logging, and failure surfaces.
  • Sharpening terminal-native workflows with tmux, Neovim, zsh, and small automation tools.
  • Turning personal infrastructure into reusable open-source patterns.

Terminal Environment

My preferred interface is a shell session with sharp tools and low friction.

editor neovim
shell zsh
session tmux
workflow git + make + docker + scripts
style keyboard-first, observable, reproducible

I care about fast feedback loops: commands that explain themselves, logs that point to causes, scripts that can be rerun safely, and local environments that make debugging boring in the best way.


Architecture Mindset

Good systems are not just systems that work once. They are systems that can be understood later.

I optimize for:

  • clear data contracts and explicit pipeline boundaries
  • scalable designs that respect failure modes
  • observability before guesswork
  • automation for repeatable operations
  • maintainable code over clever code
  • small tools that remove daily friction

Open Source Direction

I use GitHub as a public engineering notebook: experiments, infrastructure patterns, data systems, terminal tools, and the working surface of my Linux environment.

The thread across the repositories is simple: build systems that can be run, inspected, and improved by another engineer.


Top languages

LinkedIn | Email | Repositories

Pinned Loading

  1. .dotfiles.dotfilesPublic

    Neste repositorio guardo as principais configuracoes de minhas ferramentas para continuar meu fluxo de trabalho em maquinas diferentes. Mais do que isto, este repositorio serve tambem como backup.

    Shell

  2. pyspark-medallion-platformpyspark-medallion-platformPublic

    Production-grade PySpark Lakehouse implementing medallion architecture, Parquet optimization, incremental ETL, and distributed data processing.

    Jupyter Notebook

  3. document-quality-pipelinedocument-quality-pipelinePublic

    OCR-powered document validation pipeline for automated legibility analysis, extraction, and quality assurance workflows.

    Python 1

  4. api-to-postgres-pipelineapi-to-postgres-pipelinePublic

    Batch ETL pipeline that ingests public API data into PostgreSQL with incremental loading, retries, schema normalization, and Dockerized local execution.

    Python

  5. streaming-kafka-data-qualitystreaming-kafka-data-qualityPublic

    Real-time Kafka streaming pipeline with schema validation, dead-letter queues, observability, and data quality enforcement using Python and Docker.

    Python

  6. airflow-dbt-warehouseairflow-dbt-warehousePublic

    Modern analytics engineering stack using Airflow and dbt with bronze/silver/gold transformations, testing, and Docker Compose orchestration.

    Python