Table and Chunked Columns

Time: ~9 minutes. Tangible win: build a Table from multiple RecordBatches, inspect its chunked columns, and convert to/from pandas.

Table is what you'll actually get back from most pyarrow-level operations — reading a Parquet file, querying via Flight, etc. Almost nothing above the RecordBatch layer hands you a bare RecordBatch; it hands you a Table.

Primary source pyarrow documentation — pa.Table, ChunkedArray.

Exercise

import pyarrow as pa

schema = pa.schema([
    pa.field("id", pa.int64()),
    pa.field("value", pa.float64()),
])

batch1 = pa.record_batch([pa.array([1, 2, 3]), pa.array([1.1, 2.2, 3.3])], schema=schema)
batch2 = pa.record_batch([pa.array([4, 5]), pa.array([4.4, 5.5])], schema=schema)

table = pa.Table.from_batches([batch1, batch2])
print("num_rows:", table.num_rows)
print("num_columns:", table.num_columns)
print("chunks in 'id' column before combine:", table.column("id").num_chunks)

combined = table.combine_chunks()
print("chunks in 'id' column after combine:", combined.column("id").num_chunks)

df = table.to_pandas()
table_back = pa.Table.from_pandas(df, preserve_index=False)
print("Round-trip through pandas preserves row count:", table_back.num_rows == table.num_rows)
print("Round-trip data equal:", table.equals(table_back, check_metadata=False))

Verified output (pyarrow 25.0.1, pandas 2.2.6):

num_rows: 5
num_columns: 2
chunks in 'id' column before combine: 2
chunks in 'id' column after combine: 1
Round-trip through pandas preserves row count: True
Round-trip data equal: True
A real production footgun Appending many small batches one at a time can leave you with hundreds of tiny chunks, and the performance cost often isn't noticed until later. Watch chunk counts, not just row counts.

Retrieval check

You call table.column("id") and try to call an Array-only method directly on the result. What goes wrong?

Correct. Table columns are ChunkedArrays. Conflating them with plain Arrays is a common source of confusing method errors.

Not quite. A Table's column is a ChunkedArray (a sequence of Array chunks), not a bare Array — that distinction matters for which methods are available directly.

Does Table.from_batches([...]) copy the underlying array data when stitching batches together?

Correct. from_batches stitches batches into one Table's ChunkedArrays by reference, no copy. combine_chunks() is the one that copies.

No. Stitching batches into a Table is reference-only; copying only happens if you explicitly call .combine_chunks().

Practice

  1. Build a Table from 3+ RecordBatches of varying row counts.
  2. Report the chunk count before and after combine_chunks().
  3. Confirm a pandas round-trip preserves both row count and data (.equals(...) returns True).

New terms — Table, ChunkedArray — are in the glossary.

Ask the agent: "At roughly what chunk count does combine_chunks() start paying for itself on a real dataset?" Next lesson computes directly on Arrow data instead of round-tripping through pandas for every transformation: Arrow Compute.