Table and Chunked Columns
Table is what you'll actually get back from most pyarrow-level operations — reading a Parquet file, querying via Flight, etc. Almost nothing above the RecordBatch layer hands you a bare RecordBatch; it hands you a Table.
- A Table = one Schema + one
ChunkedArrayper column, where each ChunkedArray is a sequence of Array chunks (often, one chunk per source RecordBatch). Table.from_batches([...])stitches multiple same-schema RecordBatches into one Table without copying the underlying arrays — the chunks just get referenced..combine_chunks()merges a Table's multi-chunk columns into single-chunk ChunkedArrays. This does copy, but it can pay off later because some operations are faster or simpler over a single contiguous chunk.to_pandas()/Table.from_pandas()are the standard interop path with the pandas ecosystem.
pa.Table, ChunkedArray.
Exercise
import pyarrow as pa
schema = pa.schema([
pa.field("id", pa.int64()),
pa.field("value", pa.float64()),
])
batch1 = pa.record_batch([pa.array([1, 2, 3]), pa.array([1.1, 2.2, 3.3])], schema=schema)
batch2 = pa.record_batch([pa.array([4, 5]), pa.array([4.4, 5.5])], schema=schema)
table = pa.Table.from_batches([batch1, batch2])
print("num_rows:", table.num_rows)
print("num_columns:", table.num_columns)
print("chunks in 'id' column before combine:", table.column("id").num_chunks)
combined = table.combine_chunks()
print("chunks in 'id' column after combine:", combined.column("id").num_chunks)
df = table.to_pandas()
table_back = pa.Table.from_pandas(df, preserve_index=False)
print("Round-trip through pandas preserves row count:", table_back.num_rows == table.num_rows)
print("Round-trip data equal:", table.equals(table_back, check_metadata=False))
num_rows: 5
num_columns: 2
chunks in 'id' column before combine: 2
chunks in 'id' column after combine: 1
Round-trip through pandas preserves row count: True
Round-trip data equal: True
Retrieval check
You call table.column("id") and try to call an Array-only method directly on the result. What goes wrong?
Correct. Table columns are ChunkedArrays. Conflating them with plain Arrays is a common source of confusing method errors.
Not quite. A Table's column is a ChunkedArray (a sequence of Array chunks), not a bare Array — that distinction matters for which methods are available directly.
Does Table.from_batches([...]) copy the underlying array data when stitching batches together?
Correct. from_batches stitches batches into one Table's ChunkedArrays by reference, no copy. combine_chunks() is the one that copies.
No. Stitching batches into a Table is reference-only; copying only happens if you explicitly call .combine_chunks().
Practice
- Build a Table from 3+ RecordBatches of varying row counts.
- Report the chunk count before and after
combine_chunks(). - Confirm a pandas round-trip preserves both row count and data (
.equals(...)returnsTrue).
New terms — Table, ChunkedArray — are in the glossary.
combine_chunks() start paying for itself on a real dataset?" Next lesson computes directly on Arrow data instead of round-tripping through pandas for every transformation: Arrow Compute.