Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Course goals, alignment with participant profiles, and success metrics
- High-level migration approaches and risk considerations
- Setting up workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Key differences between SMP and MPP and their implications for migration
- Medallion (Bronze→Silver→Gold) design and an overview of Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Hands-on migration of a sample stored procedure to a notebook
- Mapping temp tables and cursors to DataFrame transformations
- Validation and comparison against original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM; validating results
- Measuring read/write performance improvements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL features: window functions, higher-order functions, and JSON/array handling
- Reading the Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Query tuning patterns: broadcast joins, hints, caching, and spill reduction
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactor a heavy SQL process into optimised Spark SQL
- Utilise Spark UI traces to identify and fix skew and shuffle issues
- Benchmark before/after and document tuning steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Transforming loops and cursors into vectorised DataFrame operations
- Modularisation, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactor a procedural ETL script into modular PySpark notebooks
- Introduce parametrization, unit-style tests, and reusable functions
- Code review and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assemble a Bronze→Silver→Gold pipeline orchestrated with Workflows
- Implement logging, auditing, retries, and automated validations
- Run the full pipeline, validate outputs, and prepare deployment notes
Operationalisation, Governance, and Production Readiness
- Unity Catalog governance, lineage, and access control best practices
- Cost, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook creation
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and lessons learned
- Gap analysis, recommended follow-up activities, and handover of training materials
- References, further learning paths, and support options
Requirements
- A solid understanding of data engineering principles
- Hands-on experience with SQL and stored procedures (Synapse / SQL Server)
- Familiarity with ETL orchestration concepts (ADF or equivalent tools)
Target Audience
- Technology managers possessing a data engineering background
- Data engineers seeking to transition procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption