Arrow Flight SQL

Time: ~8 minutes. Tangible win: explain what Flight SQL adds on top of plain Flight, and map its SQL-specific concepts back onto the Flight primitives from Lesson 8.

Reading-only lesson This is where Arrow shows up as an actual wire protocol competing with JDBC/ODBC — relevant if any client or internal system exposes an Arrow-native SQL interface (Dremio, DuckDB, InfluxDB 3, and several cloud warehouses support Flight SQL). There's no bundled server to stand up for this lesson, so the exercise is a close reading rather than a script.
Primary source Read the pyarrow.flight docs section covering Flight SQL command types, or the Arrow Flight SQL protocol docs for the primary source.

Exercise (written, not scripted)

Write 4–6 sentences mapping these Flight SQL concepts to the plain-Flight primitives from Lesson 8:

  1. Where does a CommandStatementQuery actually get transmitted?
  2. Once the server has planned the query, how does the client actually pull the rows?
  3. What's the practical difference between calling a Flight SQL server and calling the plain demo server you wrote in Lesson 8?

Evidence of competence

A correct answer identifies that the SQL command is sent as the command payload of a FlightDescriptor (typically via GetFlightInfo, which returns a Ticket), and that the rows still arrive via a standard do_get/DoGet stream of RecordBatches — i.e., Flight SQL changes how you ask for data, not how the data physically travels once the server agrees to send it.

Common failure modes Assuming Flight SQL replaces Flight as a separate protocol — it's a convention/schema built on Flight's existing RPCs, not a new transport. Also: assuming an ADBC driver behaves like a traditional ODBC/JDBC driver internally (row-at-a-time cursor semantics) — it's built to hand you Arrow batches, and code written assuming row-cursor semantics will underuse it.

Retrieval check

Is Arrow Flight SQL a separate transport protocol from plain Arrow Flight?

Correct. Flight SQL defines standard commands (CommandStatementQuery, etc.) that ride on Flight's existing do_get/GetFlightInfo machinery — it's a convention, not a new transport.

Not quite. Flight SQL is layered on Flight, not a separate protocol — the transport and streaming mechanics are identical to plain Flight.

Once a Flight SQL server has planned a CommandStatementQuery, how does the client actually retrieve the resulting rows?

Correct. Flight SQL changes how you ask for data (a structured command in the FlightDescriptor); the data itself still arrives via the standard do_get stream of RecordBatches.

No. There's no special retrieval RPC — rows come back through the same do_get stream of RecordBatches that plain Flight uses.

How does an ADBC driver typically hand back query results, compared to a traditional JDBC/ODBC driver?

Correct. ADBC is built around Flight SQL's batch-oriented model — code assuming row-cursor semantics will underuse it.

No. ADBC drivers hand back Arrow batches, not row-at-a-time cursor fetches like traditional ODBC/JDBC.

New terms — Flight SQL — are in the glossary.

Ask the agent: "Which real systems (Dremio, DuckDB, InfluxDB 3, cloud warehouses) actually expose Flight SQL today, and what would connecting to one from Python look like with ADBC?" That's the natural next step once you have a real server to point at — along with pyarrow.dataset for partitioned multi-file Parquet reads with pushdown across many files at once, or an end-to-end pipeline: Parquet on disk → Table → Compute filter/aggregate → served to a remote client via Flight, using only what this course covered.