Connecting a Query Engine (Spark)
The previous lesson proved the catalog works in isolation. This lesson proves the actual point of Lesson 1: an engine you didn't write any Polaris-specific code for can read and write the same catalog using a standard Iceberg Spark connector.
spark.sql.catalog.<name>.* property used below.
What the config actually says
- Spark needs the
iceberg-spark-runtimepackage for your Spark/Scala version (this run:iceberg-spark-runtime-3.5_2.12:1.5.2against Spark 3.5.2) plus the Iceberg SQL extensions enabled viaspark.sql.extensions. - The catalog is registered as
spark.sql.catalog.<name>, set toorg.apache.iceberg.spark.SparkCatalogwithcatalog-impl=org.apache.iceberg.rest.RESTCatalog— this is what tells Spark "talk Iceberg REST to this URI" instead of using a Hive Metastore or Spark's built-in catalog. - Credentials travel as
spark.sql.catalog.<name>.credentialinclientId:clientSecretform, withtoken-refresh-enabled=trueso Spark re-authenticates instead of failing when the token expires mid-session. - Once configured,
<name>.<namespace>.<table>in SQL resolves through Polaris exactly likecoursecatalog/course_dbdid viacurl— same entities, different client.
Exercise
import pyspark
from pyspark.sql import SparkSession
conf = (
pyspark.SparkConf()
.setAppName("polaris_demo")
.set("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions")
.set("spark.sql.catalog.polaris", "org.apache.iceberg.spark.SparkCatalog")
.set("spark.sql.catalog.polaris.warehouse", "coursecatalog")
.set("spark.sql.catalog.polaris.catalog-impl", "org.apache.iceberg.rest.RESTCatalog")
.set("spark.sql.catalog.polaris.uri", "http://polaris:8181/api/catalog")
.set("spark.sql.catalog.polaris.credential", "<clientId>:<clientSecret>")
.set("spark.sql.catalog.polaris.scope", "PRINCIPAL_ROLE:ALL")
.set("spark.sql.catalog.polaris.token-refresh-enabled", "true")
)
spark = SparkSession.builder.config(conf=conf).getOrCreate()
spark.sql("CREATE NAMESPACE IF NOT EXISTS polaris.course_db")
spark.sql("CREATE TABLE IF NOT EXISTS polaris.course_db.people (name STRING, role STRING) USING iceberg")
spark.sql("INSERT INTO polaris.course_db.people VALUES ('Bilbo', 'engineer'), ('Gandalf', 'staff')")
spark.sql("SELECT * FROM polaris.course_db.people ORDER BY name").show()
SELECT RESULT:
+-------+--------+
| name| role|
+-------+--------+
| Bilbo|engineer|
|Gandalf| staff|
+-------+--------+
SNAPSHOT COUNT:
+--------------+
|snapshot_count|
+--------------+
| 1|
+--------------+
snapshot_count is 1, not 2, because Iceberg's CREATE TABLE doesn't produce a snapshot by itself (no data yet) — the single multi-row INSERT is the first and only commit.
Confirming the round trip: the same table is now visible from the previous lesson's raw REST API, even though it was only ever created from Spark —
curl -s "http://localhost:8181/api/catalog/v1/coursecatalog/namespaces/course_db/tables" -H "Authorization: Bearer $TOKEN"
# {"identifiers":[{"namespace":["course_db"],"name":"people"}],"next-page-token":null}
That's Lesson 1's thesis made concrete: Spark never talked to curl, curl never talked to Spark, and they still agree perfectly on what tables exist — because both went through the same Polaris REST catalog.
getOrCreate()
On a containerized JVM under emulation (e.g. an amd64 image via Rosetta on Apple Silicon), a silent hang with near-zero CPU right after SparkSession.builder.getOrCreate() is classic entropy-starved SecureRandom initialization — not a Polaris or network problem. Fix with --conf spark.driver.extraJavaOptions=-Djava.security.egd=file:/dev/./urandom. This is exactly what happened verifying this exercise: the first attempt hung silently for 11+ minutes until this flag was added.
Retrieval check
What does catalog-impl=org.apache.iceberg.rest.RESTCatalog actually tell Spark?
Correct. This is the line that redirects Spark's catalog operations to Polaris's REST endpoint instead of a Hive Metastore or Spark's default catalog.
Not quite. catalog-impl selects the REST protocol as the transport to Polaris — it doesn't add caching or switch to an in-memory catalog.
A table created with plain CREATE TABLE (no USING iceberg) behaves nothing like the rest of this course. Why?
Correct. Omitting USING iceberg doesn't error — it silently produces a non-Iceberg table that won't show up through the Iceberg REST endpoints used everywhere else in this course.
No. Polaris doesn't reject the call — Spark just defaults to a different table format, which is what makes this mistake easy to miss.
Practice
- Run this against your own Polaris catalog with a table of your own design (different columns, different rows).
- Confirm the table is visible from the raw REST API too, even though you only created it from Spark.
- If running Spark and Polaris as separate containers, confirm you're using the Docker service name (e.g.
polaris) in the URI, notlocalhost— inside the Spark container,localhostmeans the Spark container itself.