What Polaris Is and the Problem It Solves

Time: ~7 minutes. Tangible win: explain what problem Polaris solves and for whom — every later lesson is a mechanical detail on top of this idea.

Apache Polaris is a catalog service: a system of record for where tables live and what they look like. It is not a query engine and not a storage system — it doesn't execute SQL and it doesn't store table data itself.

Primary source polaris.apache.org — project homepage covering Polaris's positioning as an Iceberg REST catalog and lakehouse catalog.

The problem: one catalog, many engines

Polaris implements the Iceberg REST Catalog specification — a vendor-neutral HTTP API for catalog operations (list namespaces, create a table, commit a snapshot). Any engine that speaks that spec — Spark, Trino, Flink, Snowflake, DuckDB, StarRocks, Doris — can read and write the same Iceberg tables through the same catalog, without a bespoke connector per engine-pair.

Before REST catalogs, each engine's Iceberg integration typically leaned on engine-specific catalog implementations (Hive Metastore, AWS Glue, engine-embedded catalogs). Getting five engines to safely share one set of tables meant either standardizing on one engine's catalog or accepting drift between them. A REST catalog gives every engine the same source of truth over HTTP.

Where Polaris comes from Polaris was donated by Snowflake to the Apache Software Foundation and is meant to be run as an independent, engine-agnostic service — you stand it up once, then point every engine at it.

Exercise: prove the service is alive and speaking Iceberg REST

Start (or confirm running) a local Polaris server and hit its OAuth token endpoint — no table operations, just proof the service is up and speaking the protocol:

curl -s -i http://localhost:8181/api/catalog/v1/oauth/tokens \
  -d 'grant_type=client_credentials&client_id=<your-client-id>&client_secret=<your-client-secret>&scope=PRINCIPAL_ROLE:ALL'

Verified output from a real local run. Exact hex IDs, ports, and timestamps will differ on your machine — the shape and status code will not.

HTTP/1.1 200 OK
Content-Type: application/json

{"access_token":"principal:root;password:...;realm:default-realm;role:ALL","scope":"PRINCIPAL_ROLE:ALL","token_type":"bearer","expires_in":3600}
Don't confuse the spec with the implementation The Iceberg REST Catalog spec is a protocol, implementable by anyone. Polaris is one specific, Apache-governed implementation of that protocol. Other REST catalog implementations exist (Tabular's, Nessie, various vendor-hosted ones) — "Polaris" and "Iceberg REST catalog" are not synonyms.

Retrieval check

Which best describes what Apache Polaris actually is?

Correct. Polaris tracks where tables live and what they look like — it doesn't execute queries and doesn't store the underlying Parquet/Avro files itself.

Not quite. Polaris is neither a query engine nor a storage system — it's the catalog layer that sits between engines and wherever table data actually lives.

Why does a REST-based catalog change the multi-engine story compared to engine-specific catalogs?

Correct. The payoff is interoperability, not performance — five engines can safely share one set of tables without standardizing on one engine's native catalog.

No. The REST catalog's value is a shared source of truth across engines, not query speed or replacing storage.

Practice

  1. Run the token exercise yourself against your own local Polaris server (or the Quickstart from RESOURCES.md if you don't have one running).
  2. In your own words (2–3 sentences), explain to someone who has only used a Hive Metastore why a REST-based catalog changes the multi-engine story.
  3. Name one engine other than Spark you'd want to point at the same Polaris catalog, and why.
Ask the agent: "What actually happens on the wire when Spark and Trino both read the same Iceberg table through Polaris — do they ever talk to each other?" Next lesson maps the entities Polaris organizes everything around: Catalog, Namespace, Table, and the principal/role objects layered on top.