RecordBatch and Arrow IPC

Time: ~9 minutes. Tangible win: build a RecordBatch from a Schema and a matching set of Arrays, and round-trip it through Arrow's IPC stream format.

RecordBatch — not Array, not Table — is the unit that actually gets sent across Arrow IPC and Arrow Flight (Lesson 8). Understanding it is a prerequisite for understanding how Arrow moves data between processes.

Primary source pyarrow IPC documentation — pa.ipc.new_stream / pa.ipc.open_stream used below.

Exercise

import pyarrow as pa

schema = pa.schema([
    pa.field("id", pa.int64()),
    pa.field("value", pa.float64()),
])

batch = pa.record_batch(
    [pa.array([1, 2, 3]), pa.array([1.1, 2.2, 3.3])],
    schema=schema,
)
print("num_rows:", batch.num_rows)
print("schema:\n", batch.schema)

sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, batch.schema) as writer:
    writer.write_batch(batch)

buf = sink.getvalue()
reader = pa.ipc.open_stream(buf)
roundtrip_batch = reader.read_next_batch()

print("Round-trip equal:", batch.equals(roundtrip_batch))

try:
    pa.record_batch([pa.array([1, 2, 3]), pa.array([1.1, 2.2])], schema=schema)
except Exception as e:
    print("Mismatched-length error (expected):", type(e).__name__)

Verified output (pyarrow 25.0.1):

num_rows: 3
schema:
 id: int64
value: double
Round-trip equal: True
Mismatched-length error (expected): ValueError
The error is the feature Trying to build a RecordBatch from arrays of different lengths raises immediately — don't work around it, expect it. The fixed-length-per-batch invariant is intentional, not a bug.

Retrieval check

You build a RecordBatch with a 3-element id array and a 2-element value array against a 2-field schema. What happens?

Correct. Mismatched array lengths are a hard error at construction time — a RecordBatch cannot exist in an inconsistent state.

Not quite. Arrow doesn't pad or silently truncate — it raises a ValueError so the batch never exists in a bad state.

What's the actual difference between a RecordBatch and a Table, given both "hold columns"?

Correct. RecordBatch is a single fixed-row-count chunk by definition; Table is the higher-level container spanning multiple batches (next lesson).

No. RecordBatch and Table are distinct: a RecordBatch is one fixed-size chunk, while a Table can span many RecordBatches.

Practice

  1. Build a RecordBatch with at least 3 columns of different types.
  2. Round-trip it through IPC and confirm .equals(...) returns True.
  3. Deliberately pass one array of the wrong length and capture the ValueError you get.

New terms — RecordBatch, IPC — are in the glossary.

Ask the agent: "What's inside the IPC stream format's schema message versus its record batch message?" Next lesson works with the higher-level container that spans multiple RecordBatches: Table.