Arrow Flight SQL
- Flight SQL is a protocol layered on top of Flight, not a replacement for it. It defines a standard set of commands (
CommandStatementQuery,CommandPreparedStatementQuery,CommandGetTables,CommandGetSchemas, ...) so that any Flight SQL client can talk to any Flight SQL server generically — no server-specific driver needed. - Those commands travel as the opaque
commandbytes inside aFlightDescriptor— the same descriptor type from Lesson 8, just carrying a structured SQL command instead of a plain path. - The actual query results still come back through an ordinary
do_getcall as a stream of Arrow RecordBatches — not row-oriented result sets like a JDBCResultSet. - ADBC (Arrow Database Connectivity) is the client API family built around this model: instead of handing you rows one cursor-fetch at a time, ADBC drivers hand you Arrow batches directly.
pyarrow.flight docs section covering Flight SQL command types, or the Arrow Flight SQL protocol docs for the primary source.
Exercise (written, not scripted)
Write 4–6 sentences mapping these Flight SQL concepts to the plain-Flight primitives from Lesson 8:
- Where does a
CommandStatementQueryactually get transmitted? - Once the server has planned the query, how does the client actually pull the rows?
- What's the practical difference between calling a Flight SQL server and calling the plain demo server you wrote in Lesson 8?
Evidence of competence
A correct answer identifies that the SQL command is sent as the command payload of a FlightDescriptor (typically via GetFlightInfo, which returns a Ticket), and that the rows still arrive via a standard do_get/DoGet stream of RecordBatches — i.e., Flight SQL changes how you ask for data, not how the data physically travels once the server agrees to send it.
Retrieval check
Is Arrow Flight SQL a separate transport protocol from plain Arrow Flight?
Correct. Flight SQL defines standard commands (CommandStatementQuery, etc.) that ride on Flight's existing do_get/GetFlightInfo machinery — it's a convention, not a new transport.
Not quite. Flight SQL is layered on Flight, not a separate protocol — the transport and streaming mechanics are identical to plain Flight.
Once a Flight SQL server has planned a CommandStatementQuery, how does the client actually retrieve the resulting rows?
Correct. Flight SQL changes how you ask for data (a structured command in the FlightDescriptor); the data itself still arrives via the standard do_get stream of RecordBatches.
No. There's no special retrieval RPC — rows come back through the same do_get stream of RecordBatches that plain Flight uses.
How does an ADBC driver typically hand back query results, compared to a traditional JDBC/ODBC driver?
Correct. ADBC is built around Flight SQL's batch-oriented model — code assuming row-cursor semantics will underuse it.
No. ADBC drivers hand back Arrow batches, not row-at-a-time cursor fetches like traditional ODBC/JDBC.
New terms — Flight SQL — are in the glossary.
pyarrow.dataset for partitioned multi-file Parquet reads with pushdown across many files at once, or an end-to-end pipeline: Parquet on disk → Table → Compute filter/aggregate → served to a remote client via Flight, using only what this course covered.