Connecting a Query Engine (Spark)

Time: ~9 minutes. Tangible win: create a namespace, create a table, insert rows, and query them entirely through Spark SQL, with Spark talking Iceberg REST underneath — no Polaris-specific code required.

The previous lesson proved the catalog works in isolation. This lesson proves the actual point of Lesson 1: an engine you didn't write any Polaris-specific code for can read and write the same catalog using a standard Iceberg Spark connector.

Primary source Iceberg Spark integration docs — reference for every spark.sql.catalog.<name>.* property used below.

What the config actually says

Exercise

import pyspark
from pyspark.sql import SparkSession

conf = (
    pyspark.SparkConf()
    .setAppName("polaris_demo")
    .set("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions")
    .set("spark.sql.catalog.polaris", "org.apache.iceberg.spark.SparkCatalog")
    .set("spark.sql.catalog.polaris.warehouse", "coursecatalog")
    .set("spark.sql.catalog.polaris.catalog-impl", "org.apache.iceberg.rest.RESTCatalog")
    .set("spark.sql.catalog.polaris.uri", "http://polaris:8181/api/catalog")
    .set("spark.sql.catalog.polaris.credential", "<clientId>:<clientSecret>")
    .set("spark.sql.catalog.polaris.scope", "PRINCIPAL_ROLE:ALL")
    .set("spark.sql.catalog.polaris.token-refresh-enabled", "true")
)
spark = SparkSession.builder.config(conf=conf).getOrCreate()

spark.sql("CREATE NAMESPACE IF NOT EXISTS polaris.course_db")
spark.sql("CREATE TABLE IF NOT EXISTS polaris.course_db.people (name STRING, role STRING) USING iceberg")
spark.sql("INSERT INTO polaris.course_db.people VALUES ('Bilbo', 'engineer'), ('Gandalf', 'staff')")
spark.sql("SELECT * FROM polaris.course_db.people ORDER BY name").show()

Verified output:

SELECT RESULT:
+-------+--------+
|   name|    role|
+-------+--------+
|  Bilbo|engineer|
|Gandalf|   staff|
+-------+--------+

SNAPSHOT COUNT:
+--------------+
|snapshot_count|
+--------------+
|             1|
+--------------+

snapshot_count is 1, not 2, because Iceberg's CREATE TABLE doesn't produce a snapshot by itself (no data yet) — the single multi-row INSERT is the first and only commit.

Confirming the round trip: the same table is now visible from the previous lesson's raw REST API, even though it was only ever created from Spark —

curl -s "http://localhost:8181/api/catalog/v1/coursecatalog/namespaces/course_db/tables" -H "Authorization: Bearer $TOKEN"
# {"identifiers":[{"namespace":["course_db"],"name":"people"}],"next-page-token":null}

That's Lesson 1's thesis made concrete: Spark never talked to curl, curl never talked to Spark, and they still agree perfectly on what tables exist — because both went through the same Polaris REST catalog.

If Spark hangs silently for minutes after getOrCreate() On a containerized JVM under emulation (e.g. an amd64 image via Rosetta on Apple Silicon), a silent hang with near-zero CPU right after SparkSession.builder.getOrCreate() is classic entropy-starved SecureRandom initialization — not a Polaris or network problem. Fix with --conf spark.driver.extraJavaOptions=-Djava.security.egd=file:/dev/./urandom. This is exactly what happened verifying this exercise: the first attempt hung silently for 11+ minutes until this flag was added.

Retrieval check

What does catalog-impl=org.apache.iceberg.rest.RESTCatalog actually tell Spark?

Correct. This is the line that redirects Spark's catalog operations to Polaris's REST endpoint instead of a Hive Metastore or Spark's default catalog.

Not quite. catalog-impl selects the REST protocol as the transport to Polaris — it doesn't add caching or switch to an in-memory catalog.

A table created with plain CREATE TABLE (no USING iceberg) behaves nothing like the rest of this course. Why?

Correct. Omitting USING iceberg doesn't error — it silently produces a non-Iceberg table that won't show up through the Iceberg REST endpoints used everywhere else in this course.

No. Polaris doesn't reject the call — Spark just defaults to a different table format, which is what makes this mistake easy to miss.

Practice

  1. Run this against your own Polaris catalog with a table of your own design (different columns, different rows).
  2. Confirm the table is visible from the raw REST API too, even though you only created it from Spark.
  3. If running Spark and Polaris as separate containers, confirm you're using the Docker service name (e.g. polaris) in the URI, not localhost — inside the Spark container, localhost means the Spark container itself.
Ask the agent: "What would change in this config if I wanted a second, independently-scoped Spark session reading the same catalog with fewer privileges?" Next lesson builds exactly that: a real RBAC chain and a real denial.