Parquet Interop
Parquet is Arrow's most common at-rest partner format, specifically because their type systems are close to 1:1 — that's why this interop feels "free" instead of requiring a translation layer.
pyarrow.parquet(pq) provideswrite_table/read_table, operating on the same Table object from earlier lessons.- Column projection (
columns=[...]) is not just "select these columns after reading" — Parquet's on-disk layout is columnar too, so unrequested columns are never read off disk at all. - Parquet files are split into row groups, each carrying its own statistics (min/max per column). A
filters=predicate can let the reader skip entire row groups whose stats prove no row could match — this is pushdown, distinct from reading everything and filtering in memory afterward. - Reading
ParquetFile(path).schema_arrowor.metadatacosts effectively nothing — it doesn't decode any row data.
Exercise
import pyarrow as pa
import pyarrow.parquet as pq
table = pa.table({
"id": [1, 2, 3, 4, 5],
"region": ["east", "east", "west", "west", "west"],
"amount": [100, 150, 200, 175, 50],
})
path = "/tmp/arrow_course_demo.parquet"
pq.write_table(table, path)
pf = pq.ParquetFile(path)
print("Row groups:", pf.metadata.num_row_groups)
print("Schema from metadata only (no data read):")
print(pf.schema_arrow)
projected = pq.read_table(path, columns=["id", "amount"])
print("Projected columns:", projected.column_names)
filtered = pq.read_table(path, filters=[("region", "=", "west")])
print("Filtered rows (region=west):", filtered.num_rows)
print(filtered.to_pydict())
Row groups: 1
Schema from metadata only (no data read):
id: int64
region: string
amount: int64
Projected columns: ['id', 'amount']
Filtered rows (region=west): 3
{'id': [3, 4, 5], 'region': ['west', 'west', 'west'], 'amount': [200, 175, 50]}
Retrieval check
You call pq.read_table(path, columns=["id", "amount"]) on a 10-column Parquet file. What happens to the other 8 columns?
Correct. Column projection skips unrequested columns at the disk-read level, not just after decoding.
Not quite. Because Parquet is columnar on disk, unrequested columns are never read at all — projection isn't a post-read filter.
Does every filters= predicate get pushed down to skip entire row groups?
Correct. Pushdown depends on the predicate being evaluable against stored min/max stats — not every predicate qualifies.
No. Some predicates do get pushed down via row-group statistics, but not universally — it depends on whether the predicate can be checked against those stats.
Practice
- Write a Parquet file with 4+ columns.
- Show that
columns=[...]returns exactly the requested subset. - Show a
filters=result matching what you'd get by filtering in pandas. - State in one sentence why the schema-only read didn't need to decode any rows.
New terms — Parquet — are in the glossary.