Databricks framework to validate Data Quality of pySpark DataFrames and Tables
-
Updated
Sep 25, 2026 - Python
Databricks framework to validate Data Quality of pySpark DataFrames and Tables
Open-source community study guide for all six Databricks certifications (Data Engineer Associate / Professional, Data Analyst Associate, ML Associate / Professional, GenAI Engineer Associate). Aligned to the 2025-2026 official exam guides. Obsidian-flavoured Markdown; PRs welcome.
Metadata-driven framework for Databricks Spark Declarative Pipelines. Config-driven, pattern based approach to batch & streaming across the medallion architecture. Deploys via Declarative Automation Bundles. Built for simplicity, extensibility, and alignment with the Databricks product roadmap.
Medallion Architecture for Data Engineering projects
Medallion analytics for the Ceres Open Data Index on Databricks — Lakeflow Declarative Pipelines (Bronze → Silver → Gold), runs on Free Edition.
Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
Bring your Claude Code skill unchanged and run it as a governed Databricks job. Publish once to a Unity Catalog volume, reuse from any job, chain skills into a pipeline (markdown in, branded PowerPoint out). No external API key. Runs on Free Edition.
Databricks SQL in Action — End-to-end medallion architecture lab using Unity Catalog, Volumes, Streaming Tables, Materialized Views, AI SQL functions, dashboards, lineage, and workflow orchestration.
Built an end-to-end retail lakehouse using the Medallion architecture with batch and streaming pipelines for scalable business analytics.
Hands-on Azure Databricks learning project — medallion pipeline, Lakeflow SDP, Delta Lake, Unity Catalog security, SQL analytics and Jobs. Built end-to-end on Azure with real executed notebooks.
Demo of Databricks Lakeflow Jobs Automation with StackQL and Databricks Asset Bundles
Obsidian-ready Databricks 80/20 learning vault with hands-on SQL, PySpark, Delta Lake, Unity Catalog, Lakeflow, MLflow, tips and labs.
End-to-end NYC Taxi data engineering pipeline using Azure Databricks, Apache Spark, Delta Lake, ADLS Gen2, Terraform, and Medallion Architecture
Built a real-time Azure lakehouse that streams synthetic ride-booking events from FastAPI through Event Hubs into Databricks, unifies them with historical data, and publishes SCD-managed facts and dimensions.
Metadata-driven schemaless MongoDB Debezium CDC ingestion into a Unity Catalog medallion lakehouse on Databricks Lakeflow Declarative Pipelines, using the VARIANT data type
Sample Databricks Asset Bundle: hotel daily performance KPIs with Lakeflow SDP, UC Metric Views, AI/BI dashboard (Brickstar styled), and Genie NL→SQL
Tsuga Logs community connector for Databricks Lakeflow Connect — scheduled log ingestion into Delta tables you own
Built an OAuth-authenticated Spark streaming lakehouse that consumes NASA GCN Fermi gamma-ray burst notices from Kafka, parses classic-text messages, and publishes a governed analytical snowflake schema.
End-to-end Azure Databricks Data Engineering Pipeline with Medallion Architecture, Delta Lake, Unity Catalog, and Lakeflow Jobs.
To associate your repository with the lakeflow topic, visit your repo's landing page and select "manage topics."