Skip to content

feat(bigquery): accelerate row-based query() with Arrow wire format - #14405

Merged
jinseopkim0 merged 1 commit into
mainfrom
feat-bigquery-arrow-query-rowbased
Sep 18, 2026
Merged

jinseopkim0 merged 1 commit into
mainfrom
feat-bigquery-arrow-query-rowbased

Conversation

@jinseopkim0

Copy link
Copy Markdown
Contributor

Enables Apache Arrow wire acceleration for the traditional BigQuery.query() API returning row-based TableResult.

Part of the BigQuery Apache Arrow support stack. Based on #14404 (page fetcher).

@jinseopkim0
jinseopkim0 added this pull request to stack #14406 September 16, 2026 17:04

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for executing fast-path queries with Arrow-formatted wire responses in BigQueryImpl, deserializing Arrow IPC schemas and record batches, and handling multi-page results using ArrowQueryPageFetcher. It also adds comprehensive unit tests covering various Arrow query scenarios. The feedback suggests optimizing the fallback path in queryRpcArrow when a query is complete but the Arrow schema is missing; instead of calling getJob and triggering redundant RPCs, the method can directly construct and return a TableResult using the standard JSON format already present in the response.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the Arrow query results format on the fast-path query execution in BigQueryImpl, introducing deserialization of Arrow IPC schemas and record batches, pagination handling via ArrowQueryPageFetcher, and corresponding unit tests. Feedback on the changes suggests extracting the duplicated logic for calculating numRows from QueryResponse in both createTableResultFromJsonResponse and queryRpcArrow into a single private helper method to improve maintainability.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for executing fast-path queries expecting Arrow-formatted wire responses in BigQueryImpl. It adds the queryRpcArrow method to handle deserialization of Arrow IPC schemas and record batches, configures pagination via ArrowQueryPageFetcher, and updates query and queryWithTimeout to route Arrow-formatted queries appropriately. Comprehensive unit tests are added in BigQueryImplTest to validate these changes. The review feedback highlights potential resource leaks in the newly added tests where BigIntVector instances are not safely closed if an exception occurs during allocation or initialization, recommending the use of try-with-resources to ensure proper cleanup.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements support for Arrow-formatted query results in the fast-path query execution path (queryRpcArrow). It refactors existing JSON response parsing into helper methods, updates queryWithTimeout to route Arrow-format queries to the new queryRpcArrow method, and adds comprehensive unit tests to verify Arrow query execution, pagination, and error handling. Feedback was provided to improve the fallback mechanism when the Arrow schema is missing or incomplete by checking if a standard JSON schema is available before throwing an exception.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the Arrow query results format in fast-path query execution by implementing queryRpcArrow, which deserializes Arrow IPC schemas and record batches, and handles pagination using ArrowQueryPageFetcher. It also updates the query method to support this fast-path execution while throwing an UnsupportedOperationException for the unsupported slow-path Arrow execution, accompanied by comprehensive unit tests. The review feedback suggests throwing an UnsupportedOperationException instead of returning an incomplete Job when jobComplete is false during Arrow query execution, ensuring consistency since the slow query path is not yet supported.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for executing fast-path queries expecting Arrow-formatted wire responses in BigQueryImpl. It introduces the queryRpcArrow method to handle deserialization of Arrow IPC schemas and record batches into standard TableResult representations, including pagination support via ArrowQueryPageFetcher. It also updates queryWithTimeout to route Arrow-format queries appropriately and adds comprehensive unit tests to validate fast-path Arrow queries, multi-page results, serialization, and error handling. There are no review comments, so I have no feedback to provide.

@jinseopkim0

Copy link
Copy Markdown
Contributor Author

@gemini-code-assist review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for executing fast-path query RPC requests expecting Arrow-formatted wire responses in BigQueryImpl. It introduces the queryRpcArrow method to deserialize Arrow IPC schemas and record batches, handles pagination using ArrowQueryPageFetcher, and integrates this flow into the main query execution path. Additionally, comprehensive unit tests have been added to BigQueryImplTest to validate various Arrow query scenarios, including fast-path execution, multi-page results, serialization, and error handling. There are no review comments to evaluate, and the implementation appears complete and well-tested.

@jinseopkim0
jinseopkim0 marked this pull request as ready for review September 18, 2026 00:55
@jinseopkim0
jinseopkim0 requested review from a team as code owners September 18, 2026 00:55
@jinseopkim0
jinseopkim0 requested a review from lqiu96 September 18, 2026 00:55
@jinseopkim0
jinseopkim0 force-pushed the feat-bigquery-arrow-query-rowbased branch from b3edf0a to 773e182 Compare September 18, 2026 01:05
@jinseopkim0
jinseopkim0 requested a review from lqiu96 September 18, 2026 18:54
Base automatically changed from feat-bigquery-arrow-page-fetcher to main September 18, 2026 19:34
@jinseopkim0
jinseopkim0 force-pushed the feat-bigquery-arrow-query-rowbased branch from ee96503 to e2c96f0 Compare September 18, 2026 19:40
@jinseopkim0
jinseopkim0 merged commit 8d12a8f into main Sep 18, 2026
210 checks passed
@jinseopkim0
jinseopkim0 deleted the feat-bigquery-arrow-query-rowbased branch September 18, 2026 20:36
jinseopkim0 added a commit that referenced this pull request Sep 18, 2026
… query() (#14409)

This PR implements slow-path execution fallback for row-based queries
requesting Arrow results format (`QueryResultsFormat.ARROW`). When
queries cannot be evaluated via the fast query path (such as queries
writing to destination tables or exceeding fast-path limits), BigQuery
job execution is triggered and table results are streamed via the
BigQuery Storage Read API in Arrow format.

Follow-up PR stacked on top of #14405.
blakeli0 pushed a commit that referenced this pull request Sep 23, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>1.92.0</summary>

##
[1.92.0](v1.91.0...v1.92.0)
(2026-09-23)


### Features

* **bigquery-jdbc:** add `EnableTimestampPicos` connection property and
its plumbing
([#14284](#14284))
([b4aa5ac](b4aa5ac))
* **bigquery-jdbc:** implement picosecond temporal math and formatting
engine
([#14286](#14286))
([2a9612a](2a9612a))
* **bigquery-jdbc:** support picosecond in REST JSON path and nested
types
([#14334](#14334))
([15ffe4a](15ffe4a))
* **bigquery-jdbc:** support picosecond in `PreparedStatement`
parameters and batching
([#14373](#14373))
([c1aac66](c1aac66))
* **bigquery-jdbc:** support picosecond timestamp in `ResultSetMetaData`
and `DatabaseMetaData`
([#14358](#14358))
([43acdd3](43acdd3))
* **bigquery-jdbc:** support picosecond timestamps in Arrow Storage Read
API and nested types
([#14332](#14332))
([b5d9aca](b5d9aca))
* **bigquery-jdbc:** support qualified project delimiter in
`DefaultDataset` property
([#14240](#14240))
([6e8d6c8](6e8d6c8))
* **bigquery:** accelerate row-based query() with Arrow wire format
([#14405](#14405))
([8d12a8f](8d12a8f))
* **bigquery:** add ArrowDeserializer helper utility
([#13943](#13943))
([d9a298b](d9a298b))
* **bigquery:** add ArrowQueryPageFetcher for Arrow query result
pagination
([#14404](#14404))
([615409f](615409f))
* **bigquery:** add ArrowQueryResult and ArrowQueryResultImpl for Arrow
result streaming
([#13944](#13944))
([a62fdf8](a62fdf8))
* **bigquery:** add Storage Read API slow-path fallback for row-based
query()
([#14409](#14409))
([26e568a](26e568a))
* **bigquery:** add zero-copy queryArrow API for Arrow VectorSchemaRoot
streaming
([#14402](#14402))
([b44ffe8](b44ffe8))
* **bigquery:** make BigQuery AutoCloseable with default no-op close
method
([#14434](#14434))
([00bf3de](00bf3de))
* **firestore:** add support for BSON types
([#13189](#13189))
([8a123d9](8a123d9))
* **gax:** add ApiCallContext and request-level settings overloads to
ResumableUploadCallable
([#14251](#14251))
([e8cbd42](e8cbd42))
* **gax:** add globalTimeout settings field to
ResumableUploadCallSettings
([#14253](#14253))
([438cda6](438cda6))
* **gax:** add resumable upload error classification and retry algorithm
([#14419](#14419))
([b70396d](b70396d))
* **gax:** add ResumableUploadCallable creation to Callables and
HttpJsonCallableFactory
([#14242](#14242))
([7de24de](7de24de))
* **gax:** implement baseline Callable and Future for resumable uploads
([#14241](#14241))
([5a54db9](5a54db9))
* **generator:** add model flag and allowlist parser for resumable
upload RPCs
([#14317](#14317))
([acc1856](acc1856))
* **generator:** emit resumable upload client surface
([#14319](#14319))
([a9fed00](a9fed00))
* **generator:** emit resumable upload settings and HttpJson upload stub
([#14321](#14321))
([c122474](c122474))
* **generator:** enable resumable upload generation for showcase
([#14325](#14325))
([f9ebd79](f9ebd79))
* **generator:** switch resumable upload specialized stubs to package
private
([#14471](#14471))
([0d4e875](0d4e875))
* **generator:** wire transport stub delegation to resumable upload
stubs
([#14322](#14322))
([cc4b980](cc4b980))
* **google/cloud/backupdr/v1beta:** add backupdr
([#14410](#14410))
([a4a47da](a4a47da))
* **google/cloud/networkservices/v1beta1:** add networkservices
([#14407](#14407))
([21c4955](21c4955))
* **pubsub:** add publish telemetry headers for publish attempt
observability
([#14338](#14338))
([c167ab8](c167ab8))
* **pubsub:** implement publish hedging to reduce tail latency
([#13735](#13735))
([b302615](b302615))
* **spanner:** Support dynamic TLS certificate and key rotation for
Spanner Omni
([#14456](#14456))
([ffc745c](ffc745c))
* **storage/control:** add delete folder recursive sample
([#13642](#13642))
([f4b1b46](f4b1b46))
* **storage/control:** add delete folder recursive sample
([#14397](#14397))
([2c01d55](2c01d55))


### Bug Fixes

* **auth:** restore transportFactory upon deserialization in
InternalAwsSecurityCredentialsSupplier
([#14340](#14340))
([beea42f](beea42f))
* **bigquery-jdbc:** ensure row ordering in PCNT IT
([#14330](#14330))
([a16f048](a16f048))
* **bigquery-jdbc:** fix htapi fallback due to permission logic
([#14418](#14418))
([21e6dc8](21e6dc8))
* **bigquery-jdbc:** fix Timestamp assertions
([#14290](#14290))
([533ba14](533ba14))
* **bigquery-jdbc:** handle null parameters in Storage Write API bulk
inserts
([#14270](#14270))
([dd2c41a](dd2c41a)),
refs
[#14066](#14066)
* **bigquery-jdbc:** handle SQL NULLs in ResultSet primitive getters
([#14383](#14383))
([8e464fe](8e464fe)),
refs
[#14371](#14371)
* **bigquery:** default Arrow pagination stream location to US instead
of global
([#14458](#14458))
([2775eb1](2775eb1))
* **bigquery:** preserve page token and paginate correctly in Arrow
query when maxResults is set
([#14469](#14469))
([f5601f4](f5601f4))
* **bigquery:** use first page row count for Arrow query pagination
offset
([#14466](#14466))
([9d10dd0](9d10dd0))
* **bigtable:** don't notify config listeners while holding the manager
lock
([#14294](#14294))
([4426ccd](4426ccd))
* **bigtable:** fall back to classic path when per-RPC CallCredentials
are set on session path
([#14477](#14477))
([57bacb0](57bacb0))
* **bigtable:** fix abnormal session closures and scale-up in session
pool
([#14431](#14431))
([6361ecd](6361ecd))
* **biqguery:** fix undeclared QueryParameter wiring in QueryStatistics
([#14401](#14401))
([64cf1d3](64cf1d3))
* **bom:** restore google-cloud-spanner-jdbc to libraries-bom
([#14362](#14362))
([bc7be5e](bc7be5e)),
refs
[#14347](#14347)
* **spanner:** honor maxAttempts and totalTimeout in streaming resume
loop
([#14370](#14370))
([305f47d](305f47d))
* **spanner:** only set snapshot isolation read timestamp for SI or
optimistic txns in CloudClientExecutor
([#14346](#14346))
([54c0d0f](54c0d0f))
* **spanner:** prevent statement cancellation race in
AbstractBaseUnitOfWork
([#14283](#14283))
([d9a8eef](d9a8eef))
* **spanner:** re-enable ITInstanceAdminTest on cloud-devel and
cloud-staging
([#14281](#14281))
([89a8268](89a8268))


### Performance Improvements

* **spanner:** stop re-parsing the request id on every RPC
([#14353](#14353))
([46108f4](46108f4))


### Documentation

* Add a Http/Json Post-Quantum Cryptography Guide
([#13963](#13963))
([fcc65b0](fcc65b0))
* **bigquery:** add QueryArrow code sample and document JDK 17+ JVM
requirements
([#14437](#14437))
([bd363f6](bd363f6))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: release-please[bot] <55107282+release-please[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants