Reference — the vocabulary of Apache Arrow's data model and ecosystem. All lessons use terms consistently with this page.
Core data model
Schema
An ordered list of Fields, plus optional schema-level metadata (a bytes-keyed dict). Order matters — two schemas with the same fields in a different order are not equal. [docs]
Field
name + type + nullable flag + optional metadata. Describes one column, not its data. Fields default to nullable=True unless set otherwise. [docs]
Array
An immutable, contiguous columnar buffer of typed values, built with pa.array(...). The actual data a Field/Schema only describes abstractly. [docs]
Buffer
A raw block of memory backing an Array (e.g. the data buffer, the validity bitmap buffer). Accessible via array.buffers(); comparing buffer addresses is how this course proves zero-copy behavior. [docs]
Validity bitmap
A separate bitmap, one bit per element, that tracks null/not-null for an Array — nulls are never encoded as a sentinel value inside the data buffer itself. [docs]
Zero-copy
Sharing the same underlying memory buffer instead of serializing/copying data — possible because Arrow's memory layout is a published, language-independent spec. Conditional, not unconditional: it holds when representations already match (e.g. float64 to NumPy), not for every type.
RecordBatch
One Schema + one Array per field, all the same length. A fixed-size, fixed-row-count chunk of columnar data — the unit that actually moves across Arrow IPC and Arrow Flight. Mismatched array lengths are a hard error at construction. [docs]
Table
One Schema + one ChunkedArray per column. The higher-level container that can span multiple RecordBatches — what most pyarrow-level operations (Parquet reads, Flight queries) actually hand you. [docs]
ChunkedArray
A Table column's actual type: a sequence of Array chunks, often one per source RecordBatch. Distinct from a plain Array — some operations need .combine_chunks() or .chunk(i) first. [docs]
Moving data
IPC (Inter-Process Communication)
Arrow's serialization format (pyarrow.ipc) for writing a stream of RecordBatches to bytes and reading them back, with no row-by-row translation. Used for files and as the wire format inside Flight. [docs]
Flight
A gRPC-based streaming protocol for moving RecordBatches between processes over a network, with no per-row (de)serialization. Not a request/response REST API. [docs]
FlightDescriptor
Identifies a dataset to a Flight server, by path or an opaque command (the latter is how Flight SQL carries a SQL statement). [docs]
Ticket
What a client hands to do_get to actually retrieve a Flight data stream, typically obtained from a prior FlightInfo/GetFlightInfo lookup. [docs]
do_get / do_put
The two core Flight data-movement RPCs: do_get streams server → client, do_put streams client → server. list_flights / get_flight_info are metadata/discovery RPCs instead. [docs]
Flight SQL
A protocol layered on top of Flight (not a replacement) defining standard commands (CommandStatementQuery, CommandGetTables, ...) so any Flight SQL client can talk to any Flight SQL server generically. Results still arrive via ordinary do_get RecordBatch streams. [docs]
ADBC (Arrow Database Connectivity)
The client API family built around Flight SQL's model — drivers hand back Arrow batches directly instead of row-at-a-time cursor fetches like traditional JDBC/ODBC. [docs]
At-rest interop
Parquet
An on-disk columnar file format Arrow interoperates with closely, because their type systems are close to 1:1. Distinct from Arrow: Parquet is at-rest and compressed; Arrow is in-memory and decoded. [docs]
Column projection
Requesting only specific columns (columns=[...]) from a Parquet read. Because Parquet's on-disk layout is columnar too, unrequested columns are never read off disk at all.
Row group
A horizontal partition of a Parquet file, each carrying its own per-column statistics (min/max). A filters= predicate that can be checked against those stats lets the reader skip entire row groups — predicate pushdown.
Compute kernel
A vectorized function in pyarrow.compute (e.g. pc.greater, pc.filter, pc.sum) that operates directly on Arrow Arrays/ChunkedArrays/Tables and always returns a new object — never mutates in place. Comparisons involving null propagate null, not False. [docs]