Building Batch Data Analytics Solutions on AWS

Learn to build batch data analytics solutions using Amazon EMR, Apache Spark, Hive, AWS Glue, and related AWS services, with a focus on ingestion, storage, processing, security, monitoring, performance, and cost management.

Course Overview

Building Batch Data Analytics Solutions on AWS is an intermediate-level AWS classroom course focused on designing and implementing batch data analytics solutions using Amazon EMR. The course covers data collection, ingestion, storage, processing, cataloging, security, monitoring, and cost management within batch analytics pipelines.

Learners explore Amazon EMR with Apache Spark, Apache Hive, and Apache HBase, along with integration with AWS Glue, AWS Lake Formation, and AWS Step Functions. The course also covers EMR Notebooks for analytics and machine learning workloads, storage optimization, cluster architecture, security, performance, and troubleshooting. Through interactive demonstrations, practice labs, and a design activity, participants apply these technologies to build and manage batch data analytics workflows. The course is intended for data platform engineers, architects, and operators who build and manage data analytics pipelines.

Course Objective

  • Compare the features and benefits of data warehouses, data lakes, and modern data architectures
  • Design and implement a batch data analytics solution
  • Identify and apply appropriate techniques, including compression, to optimize data storage
  • Select and deploy appropriate options to ingest, transform, and store data
  • Choose appropriate instance and node types, clusters, auto scaling, and network topology for specific business use cases
  • Understand how data storage and processing affect analysis and visualization mechanisms for actionable business insights
  • Secure data at rest and in transit
  • Monitor analytics workloads to identify and remediate problems
  • Apply cost management best practices

Pre-requisites

Students should have a minimum of one year of experience managing open-source data frameworks such as Apache Spark or Apache Hadoop. Recommended AWS preparation includes AWS Technical Essentials or Architecting on AWS, along with Building Data Lakes on AWS or Getting Started with AWS Glue.

Course Curriculum