Parquet Interop

Time: ~10 minutes. Tangible win: write a Table to Parquet and read it back with column projection and row-group-level filtering, and read Parquet metadata without touching the data.

Parquet is Arrow's most common at-rest partner format, specifically because their type systems are close to 1:1 — that's why this interop feels "free" instead of requiring a translation layer.

Primary source pyarrow Parquet documentation and Apache Parquet.

Exercise

import pyarrow as pa
import pyarrow.parquet as pq

table = pa.table({
    "id": [1, 2, 3, 4, 5],
    "region": ["east", "east", "west", "west", "west"],
    "amount": [100, 150, 200, 175, 50],
})

path = "/tmp/arrow_course_demo.parquet"
pq.write_table(table, path)

pf = pq.ParquetFile(path)
print("Row groups:", pf.metadata.num_row_groups)
print("Schema from metadata only (no data read):")
print(pf.schema_arrow)

projected = pq.read_table(path, columns=["id", "amount"])
print("Projected columns:", projected.column_names)

filtered = pq.read_table(path, filters=[("region", "=", "west")])
print("Filtered rows (region=west):", filtered.num_rows)
print(filtered.to_pydict())

Verified output (pyarrow 25.0.1):

Row groups: 1
Schema from metadata only (no data read):
id: int64
region: string
amount: int64
Projected columns: ['id', 'amount']
Filtered rows (region=west): 3
{'id': [3, 4, 5], 'region': ['west', 'west', 'west'], 'amount': [200, 175, 50]}
Parquet still decodes — it doesn't hand Arrow raw bytes Parquet's on-disk bytes are not literally reused as Arrow's in-memory buffers with zero decode step. Parquet uses its own encodings (dictionary, RLE, etc.) and gets decoded into Arrow buffers on read. The real efficiency story is "skips columns and row groups it doesn't need," not "shares raw bytes with Arrow."

Retrieval check

You call pq.read_table(path, columns=["id", "amount"]) on a 10-column Parquet file. What happens to the other 8 columns?

Correct. Column projection skips unrequested columns at the disk-read level, not just after decoding.

Not quite. Because Parquet is columnar on disk, unrequested columns are never read at all — projection isn't a post-read filter.

Does every filters= predicate get pushed down to skip entire row groups?

Correct. Pushdown depends on the predicate being evaluable against stored min/max stats — not every predicate qualifies.

No. Some predicates do get pushed down via row-group statistics, but not universally — it depends on whether the predicate can be checked against those stats.

Practice

  1. Write a Parquet file with 4+ columns.
  2. Show that columns=[...] returns exactly the requested subset.
  3. Show a filters= result matching what you'd get by filtering in pandas.
  4. State in one sentence why the schema-only read didn't need to decode any rows.

New terms — Parquet — are in the glossary.

Ask the agent: "Which kinds of filter predicates actually get pushed down to row-group stats, and which ones don't?" Next lesson moves data across a network between processes, not just within one process: Arrow Flight.