Starburst Essentials

Master federated SQL querying with Starburst in this comprehensive 5-day program. Learn to query,
optimize, and manage distributed data across multiple sources without data movement. From SQL
fundamentals to data mesh architecture, build enterprise-grade analytics platforms.

Course Overview

The Starburst Essentials training program is a comprehensive 5-day course designed to equip data
professionals with the skills to query, manage, and optimize distributed data using Starburst, a
fast and scalable SQL query engine designed for running large-scale data queries across distributed
systems. This course is composed to help you master all the essential objectives to effectively use
Starburst in enterprise environments. You will learn foundational concepts, federated querying across
multiple data sources, advanced SQL techniques, data lake optimization, Apache Iceberg table formats,
streaming data integration with Kafka, security and governance, and data mesh architecture. Participants
gain hands-on experience in building scalable data pipelines, implementing role-based access control,
and integrating streaming data. By the end of the training, you will understand how to design and
operate high-performance, enterprise-grade data platforms using Starburst.

Course Objective

• Differentiate components in a Starburst cluster and understand their roles in distributed query execution
• Execute federated queries by joining multiple data sources without moving data from their source
• Leverage SQL functions including transforms, aggregates, and windowing functions for complex analytics
• Apply performance optimization techniques using SQL nuances, approximation strategies, and cost-based optimization
• Construct analytical queries using rollup, cube, and windowing functions for advanced business intelligence
• Use Hive and Iceberg table formats to construct, populate, query, and modify partitioned data lake tables
• Employ file size, format, and hierarchy strategies to improve query performance on large-scale data lakes
• Create and validate role-based access control (RBAC) and attribute-based access control (ABAC) policies
• Build data engineering pipelines with Starburst Galaxy for scalable data processing
• Integrate streaming data from Confluent Kafka into Starburst for real-time analytics
• Design and implement data mesh architecture using Starburst as the federated query engine
• Understand Starburst high availability, scalability, security, and authentication mechanisms

Pre-requisites

• Intermediate experience with SQL (SELECT, WHERE, JOIN, GROUP BY statements)
• Basic understanding of databases and data storage concepts
• Familiarity with data integration tools and techniques
• Basic understanding of distributed systems and big data concepts
• Familiarity with cloud platforms (AWS, Azure, or Google Cloud Platform)
• Basic knowledge of Linux and command-line interface (CLI)

Course Curriculum