Schema and Field
Schema is the blueprint every other structure in this course (Array, RecordBatch, Table) is validated against. You can't reason about the rest of the data model without it.
- A Field is
name + type + nullable flag + optional metadata. It describes one column, not the data in it. - A Schema is an ordered list of Fields, plus optional schema-level metadata (a bytes-keyed dict, separate from field metadata).
- Arrow has its own type system (
pa.int64(),pa.string(),pa.timestamp('us'),pa.list_(...),pa.struct([...]), etc.) — these are not Python types or numpy dtypes, though they map to/from them. - Order matters for schema equality; two schemas with the same fields in a different order are not equal.
pa.schema() / pa.field() API surface used below.
Exercise
import pyarrow as pa
schema = pa.schema([
pa.field("id", pa.int64(), nullable=False),
pa.field("name", pa.string()),
pa.field("tags", pa.list_(pa.string())),
pa.field("address", pa.struct([
pa.field("city", pa.string()),
pa.field("zip", pa.string()),
])),
])
print(schema)
print("Field names:", schema.names)
print("id field nullable:", schema.field("id").nullable)
print("name field nullable (default):", schema.field("name").nullable)
reordered = pa.schema([schema.field("name"), schema.field("id")])
print("Reordered schema equal to original order:",
reordered.equals(pa.schema([schema.field("id"), schema.field("name")])))
id: int64 not null
name: string
tags: list<item: string>
child 0, item: string
address: struct<city: string, zip: string>
child 0, city: string
child 1, zip: string
Field names: ['id', 'name', 'tags', 'address']
id field nullable: False
name field nullable (default): True
Reordered schema equal to original order: False
True unless you set nullable=False explicitly — the opposite of what many newcomers assume.
Retrieval check
If you construct a field with pa.field("name", pa.string()) and don't pass nullable, what is it?
Correct. The default is nullable; you have to explicitly pass nullable=False to forbid nulls in that column.
Not quite. Arrow fields default to nullable True — the opposite of the common assumption.
Two schemas contain the exact same fields, but the "id" and "name" fields are in a different order between them. Are the schemas equal?
Correct. Order matters for schema equality, as the exercise's reordered-schema comparison shows returning False.
No. Schema equality is order-sensitive — reordering fields with otherwise identical content makes schemas compare unequal.
Where does per-column metadata (e.g. a description of just the "id" field) get attached, versus metadata about the whole dataset?
Correct. Field metadata and schema metadata are stored and inspected separately — conflating them is a common early mistake.
No. Field-level and schema-level metadata are two distinct mechanisms, attached and read separately.
Practice
- Build a schema for a dataset you actually work with: 3+ fields, at least one nested type (
list_orstruct). - Print it and state which fields are nullable — and why that's correct for your data, not just "because it's the default."
New terms — Schema, Field — are in the glossary.
pa.timestamp('us') and pa.timestamp('us', tz='UTC')." Next lesson moves from this abstract description to the concrete columnar data structure it describes: Array.