Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundations (Theory):

  • System Architecture
  • Resilient Distributed Datasets (RDDs)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Applying Core Concepts in Databricks (Hands-on Workshop):

  • Practical exercises using the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • User-Defined Functions (UDFs)
  • Introduction to the DataSet API
  • Stream processing

Exploring Deployment in AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Running example jobs in both environments
  • Evaluating the advantages and limitations of each

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming experience (ideally in Python or Scala)

Foundational knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Provisional Upcoming Courses (Require 5+ participants)

Related Categories