Skip to content

Latest commit

 

History

31 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Data Engineer | Data Analyst | Data Warehouse | ETL | Cloud Data Platform

LinkedIn Email


About Me

I've spent the past 9+ years turning messy data into something teams can actually rely on — from ETL/ELT pipelines and data warehouses to streaming systems and cloud platforms. These days I'm also exploring how AI/LLMs fit into that world, building data-driven solutions on top of it.


Data Engineering

Batch and Streaming Pipelines Lakehouse Architecture Data Quality and Observability AI and LLM Data Processing

Layer Tech Stack
Orchestration Apache Airflow, Dagster, Prefect, Apache NiFi
Language Python, SQL
Processing Apache Spark, Apache Flink
Streaming Apache Kafka
Ingestion Airbyte, Debezium, Talend, Pentaho, IBM DataStage, SSIS
Cloud Computing AWS, GCP, Azure, Alibaba Cloud
Lakehouse Apache Iceberg, Amazon S3, Google Cloud Storage, SeaweedFS
Data Platform Databricks, Microsoft Fabric, Snowflake
Transformation dbt, SQLMesh
Warehouse and Query BigQuery, Redshift, ClickHouse, DuckDB, PostgreSQL, SQL Server, Oracle, Trino
Search and Enrichment OpenSearch, MongoDB
Monitoring Grafana, Prometheus
BI and Visualization Power BI, Tableau, IBM Cognos, Metabase, Plotly Dash, Superset
DevOps and Containerization Docker, Podman
Collaboration Jira, Confluence
AI/LLM Data Processing LLM-driven data workflows

Core Skills

  • Data Engineering and Data Analysis
  • Data Warehousing with Data Vault 2.0 and Kimball methodology
  • ETL and data pipeline development
  • Business Intelligence and dashboard development
  • SQL development across PostgreSQL, Oracle, MySQL, SQL Server, BigQuery, Redshift
  • Workflow orchestration with Apache Airflow
  • Streaming and integration with Kafka and Apache NiFi
  • Cloud data solutions on Google Cloud Platform, AWS, and Alibaba Cloud
  • Data platforms including Databricks, Microsoft Fabric, and Snowflake

Projects

Use Case Projects

databricks-spark-medallion airflow-duckdb-dash dagster-iceberg-metabase prefect-seaweedfs-duckdb trino-mongodb-metabase datavault-dbt kafka-cdc-jdbc-sink kafka-cdc-flink

Learning Repos

dbt-postgres-story sql-story-postgresql

Mentoring Repos

kafka-hands-on-basic docker-hands-on-basic


Tech Stack

Data Engineering

Apache Airflow Apache Kafka Apache Spark Apache Flink Apache NiFi Dagster Apache Iceberg dbt Talend Pentaho Airbyte SQLMesh

Databases and Warehousing

PostgreSQL Oracle MySQL SQL Server MongoDB ClickHouse DuckDB BigQuery Amazon Redshift OpenSearch SeaweedFS Trino

Cloud and Data Platforms

Google Cloud AWS Microsoft Azure Alibaba Cloud Databricks Microsoft Fabric

Analytics and Programming

Python Power BI Tableau IBM Cognos Metabase Grafana Plotly Dash Superset

Tools and Observability

Docker Podman Prometheus Jira Confluence


GitHub

GitHub Repositories GitHub Profile


Connect

About

GitHub profile

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors