Arrow Compute

Time: ~9 minutes. Tangible win: use pyarrow.compute to filter, compare, and aggregate a Table without converting to pandas first — and correctly predict how nulls behave in comparisons.

A lot of pipelines do "load into Arrow → convert to pandas → transform → convert back" purely out of habit. Compute lets you filter/aggregate in Arrow directly, which is faster and skips the extra conversions. It's also required reading before Parquet predicate pushdown (next lesson) makes sense.

Primary source pyarrow.compute API reference.

Exercise

import pyarrow as pa
import pyarrow.compute as pc

table = pa.table({
    "region": ["east", "east", "west", "west", "west"],
    "amount": [100, 150, 200, None, 50],
})

mask = pc.greater(table["amount"], 90)
print("Mask:", mask.to_pylist())

filtered = table.filter(mask)
print("Filtered rows:", filtered.num_rows)
print(filtered.to_pydict())

grouped = table.group_by("region").aggregate([("amount", "sum")])
print(grouped.to_pydict())

Verified output (pyarrow 25.0.1):

Mask: [True, True, True, None, False]
Filtered rows: 3
{'region': ['east', 'east', 'west'], 'amount': [100, 150, 200]}
{'region': ['east', 'west'], 'amount_sum': [250, 250]}
Null is not False The row with amount=None produces a mask entry of None, not False or True, and .filter() drops it — it never "failed the comparison," it was simply unknown. Code that does if not mask[i]: instead of proper null-aware filtering will misclassify those rows.

Retrieval check

pc.greater(table["amount"], 90) is evaluated on a row where amount is null. What does that mask entry contain?

Correct. Comparisons involving null propagate null, not False. Treating them as equivalent misclassifies rows.

Not quite. A comparison against null produces null, distinct from both True and False.

After calling pc.greater(table["amount"], 90), is table itself modified?

Correct. Every pyarrow.compute kernel returns a new Arrow object. table is unchanged after the call.

No. pyarrow.compute kernels never mutate in place — they return new Arrow objects, leaving the input Table untouched.

Practice

  1. On your own table with at least one null in the filtered column, show the raw mask (proving it contains None, not False, for the null row).
  2. Show the filtered result.
  3. Compute one group_by(...).aggregate(...) and confirm the numbers by hand.

New terms — compute kernel — are in the glossary.

Ask the agent: "How would I write a filter that treats null amount as 'include' instead of 'exclude'?" Next lesson gets this same data on and off disk efficiently, in a format Arrow reads natively: Parquet.