Schema and Field

Time: ~8 minutes. Tangible win: construct and inspect an Arrow Schema made of typed Fields, and know what schema equality actually checks.

Schema is the blueprint every other structure in this course (Array, RecordBatch, Table) is validated against. You can't reason about the rest of the data model without it.

Primary source pyarrow documentation — the pa.schema() / pa.field() API surface used below.

Exercise

import pyarrow as pa

schema = pa.schema([
    pa.field("id", pa.int64(), nullable=False),
    pa.field("name", pa.string()),
    pa.field("tags", pa.list_(pa.string())),
    pa.field("address", pa.struct([
        pa.field("city", pa.string()),
        pa.field("zip", pa.string()),
    ])),
])

print(schema)
print("Field names:", schema.names)
print("id field nullable:", schema.field("id").nullable)
print("name field nullable (default):", schema.field("name").nullable)

reordered = pa.schema([schema.field("name"), schema.field("id")])
print("Reordered schema equal to original order:",
      reordered.equals(pa.schema([schema.field("id"), schema.field("name")])))

Verified output (pyarrow 25.0.1):

id: int64 not null
name: string
tags: list<item: string>
  child 0, item: string
address: struct<city: string, zip: string>
  child 0, city: string
  child 1, zip: string
Field names: ['id', 'name', 'tags', 'address']
id field nullable: False
name field nullable (default): True
Reordered schema equal to original order: False
Nullable defaults the way you don't expect Fields default to nullable True unless you set nullable=False explicitly — the opposite of what many newcomers assume.

Retrieval check

If you construct a field with pa.field("name", pa.string()) and don't pass nullable, what is it?

Correct. The default is nullable; you have to explicitly pass nullable=False to forbid nulls in that column.

Not quite. Arrow fields default to nullable True — the opposite of the common assumption.

Two schemas contain the exact same fields, but the "id" and "name" fields are in a different order between them. Are the schemas equal?

Correct. Order matters for schema equality, as the exercise's reordered-schema comparison shows returning False.

No. Schema equality is order-sensitive — reordering fields with otherwise identical content makes schemas compare unequal.

Where does per-column metadata (e.g. a description of just the "id" field) get attached, versus metadata about the whole dataset?

Correct. Field metadata and schema metadata are stored and inspected separately — conflating them is a common early mistake.

No. Field-level and schema-level metadata are two distinct mechanisms, attached and read separately.

Practice

  1. Build a schema for a dataset you actually work with: 3+ fields, at least one nested type (list_ or struct).
  2. Print it and state which fields are nullable — and why that's correct for your data, not just "because it's the default."

New terms — Schema, Field — are in the glossary.

Ask the agent: "Show me a schema with a timestamp field and explain the difference between pa.timestamp('us') and pa.timestamp('us', tz='UTC')." Next lesson moves from this abstract description to the concrete columnar data structure it describes: Array.