Apache Arrow Course
A hands-on course on Apache Arrow using the Python binding (pyarrow). Nine lessons, one concept each, in order — each builds on the previous one's artifacts, so work through them in sequence rather than skipping around. Every code sample was actually run and its output verified (pyarrow 25.0.1, pandas 2.2.6, numpy 2.2.6).
Foundations
- 01 What Arrow Is and Why It Exists The in-memory columnar format, and the zero-copy interop payoff that follows from a published memory layout. ~7 min
- 02 Schema and Field Construct and inspect a typed Schema, including nested types, and know what schema equality actually checks. ~8 min
Data Model
- 03 Array and the Validity Bitmap How nulls are represented without sentinel values, and why slicing is zero-copy. ~8 min
- 04 RecordBatch and Arrow IPC Build a fixed-size RecordBatch and round-trip it through Arrow's IPC stream format. ~9 min
- 05 Table and Chunked Columns Stitch RecordBatches into a Table, understand ChunkedArray, and round-trip through pandas. ~9 min
Compute & Storage
- 06 Arrow Compute Filter and aggregate a Table with pyarrow.compute, and correctly predict null propagation. ~9 min
- 07 Parquet Interop Column projection, row-group predicate pushdown, and metadata-only reads. ~10 min
Distributed Ecosystem
- 08 Arrow Flight Run a minimal Flight server and client and confirm a byte-identical round-trip over gRPC. ~10 min
- 09 Arrow Flight SQL What Flight SQL adds on top of plain Flight and ADBC's batch-oriented client model. Reading-only. ~8 min
Reference
Course: Apache Arrow Course - 9 lessons on the in-memory columnar format, its data model, and its ecosystem (Compute, Parquet, Flight, Flight SQL).
Total course time: ~78 minutes · 9 lessons · Interactive checks in each lesson