feat(bigquery): add ArrowDeserializer helper utility - #13943
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces the ArrowDeserializer class to convert Apache Arrow schemas and record batches into BigQuery schemas and FieldValueList rows. Feedback on the implementation highlights several critical issues: a potential native memory leak if vector creation fails, a bug in repeated record deserialization where subfields are omitted, missing defensive checks for schema and size mismatches, and an inefficient timestamp conversion that should cast to TimeStampVector and handle different time units.
e149afc to
796b77f
Compare
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces comprehensive documentation (including developer, user, and storage API guides) and adds support for Apache Arrow serialization and deserialization in the BigQuery client and JDBC driver. It also introduces a centralized type registry and temporal utilities for optimized type coercion. The review feedback highlights several critical improvements: optimizing memory allocation and preventing resource leaks in ArrowDeserializer by passing the BufferAllocator and using proper try-finally blocks; avoiding precision loss in timestamp formatting by using integer division; ensuring that the new Arrow configuration options are actually transmitted to the BigQuery backend by serializing them in QueryJobConfiguration.toPb(); and adhering to timezone handling rules by using legacy Calendar manipulation and LocalTime.ofNanoOfDay for java.sql.Time conversions.
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces Apache Arrow serialization options and deserialization support to the BigQuery client library, alongside extensive documentation and optimized type registry utilities for the JDBC driver. The review feedback highlights critical issues where the new Arrow configuration fields in QueryJobConfiguration are not serialized or deserialized in toPb() and fromPb(), rendering them ineffective. Additionally, the reviewer recommends reusing a single BufferAllocator in ArrowDeserializer to avoid high allocation overhead, ensuring exception-safe LIFO cleanup of FieldVector instances to prevent memory leaks, and using a timezone-aware manual conversion when converting java.sql.Time to java.time.LocalTime to preserve millisecond precision.
9056c4c to
9ddd9d2
Compare
9ddd9d2 to
712b972
Compare
f139009 to
dd8bc5a
Compare
…ull vector values
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces ArrowDeserializer and ArrowPojoUtils to deserialize Apache Arrow IPC record batches and schemas into BigQuery Veneer objects, accompanied by comprehensive unit tests. Feedback was provided to optimize performance in ArrowDeserializer.arrowVectorToFieldValue by passing individual field components directly rather than instantiating new Field and Field.Builder objects for every element of a repeated field, thereby reducing GC pressure.
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces ArrowDeserializer and ArrowPojoUtils to handle the deserialization of Apache Arrow IPC record batches and the conversion of Arrow schemas into BigQuery Veneer objects, accompanied by comprehensive unit tests. The review feedback suggests adding a validation check in ArrowPojoUtils.getSubFields to throw an IllegalArgumentException if an Arrow Struct field has no children, as BigQuery does not support empty RECORD (STRUCT) fields.
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces ArrowDeserializer and ArrowPojoUtils to handle the deserialization of Apache Arrow IPC record batches and schema definitions into BigQuery Veneer objects, accompanied by comprehensive unit tests. The review feedback suggests two key improvements in ArrowDeserializer: first, replacing structVector.size() with structVector.getChildVectors().size() to clarify that it refers to child fields rather than row count; second, optimizing performance by using specific vector primitive getters instead of vector.getObject(rowIndex) to avoid boxing and object allocation overhead.
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces ArrowDeserializer and ArrowPojoUtils to handle the deserialization of Apache Arrow record batches and schemas into BigQuery Veneer objects, complete with comprehensive unit tests. The review feedback suggests addressing boundary condition checks in loadArrowRows when maxResults is reached to prevent returning incorrect continuation signals. Additionally, it recommends using toPlainString() for BigDecimal values to avoid scientific notation formatting, ensuring compatibility with BigQuery's expected decimal format.
…mal with toPlainString
|
@gemini-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces two new internal helper classes, ArrowDeserializer and ArrowPojoUtils, along with comprehensive unit tests. ArrowDeserializer manages Arrow vector memory and deserializes IPC record batches into BigQuery Veneer objects, while ArrowPojoUtils handles the conversion of Apache Arrow POJO definitions (such as Schema, Field, and ArrowType) into BigQuery Veneer objects. There are no review comments, and we have no feedback to provide.
lqiu96
left a comment
There was a problem hiding this comment.
LGTM. Added a few comments but nothing major from what I can see.
There is a lot of conversion logic and unit tests to confirm the behavior. We may run into some issues when we run ITs against real calls, but we can always fix that when we get there.
🤖 I have created a release *beep* *boop* --- <details><summary>1.92.0</summary> ## [1.92.0](v1.91.0...v1.92.0) (2026-09-23) ### Features * **bigquery-jdbc:** add `EnableTimestampPicos` connection property and its plumbing ([#14284](#14284)) ([b4aa5ac](b4aa5ac)) * **bigquery-jdbc:** implement picosecond temporal math and formatting engine ([#14286](#14286)) ([2a9612a](2a9612a)) * **bigquery-jdbc:** support picosecond in REST JSON path and nested types ([#14334](#14334)) ([15ffe4a](15ffe4a)) * **bigquery-jdbc:** support picosecond in `PreparedStatement` parameters and batching ([#14373](#14373)) ([c1aac66](c1aac66)) * **bigquery-jdbc:** support picosecond timestamp in `ResultSetMetaData` and `DatabaseMetaData` ([#14358](#14358)) ([43acdd3](43acdd3)) * **bigquery-jdbc:** support picosecond timestamps in Arrow Storage Read API and nested types ([#14332](#14332)) ([b5d9aca](b5d9aca)) * **bigquery-jdbc:** support qualified project delimiter in `DefaultDataset` property ([#14240](#14240)) ([6e8d6c8](6e8d6c8)) * **bigquery:** accelerate row-based query() with Arrow wire format ([#14405](#14405)) ([8d12a8f](8d12a8f)) * **bigquery:** add ArrowDeserializer helper utility ([#13943](#13943)) ([d9a298b](d9a298b)) * **bigquery:** add ArrowQueryPageFetcher for Arrow query result pagination ([#14404](#14404)) ([615409f](615409f)) * **bigquery:** add ArrowQueryResult and ArrowQueryResultImpl for Arrow result streaming ([#13944](#13944)) ([a62fdf8](a62fdf8)) * **bigquery:** add Storage Read API slow-path fallback for row-based query() ([#14409](#14409)) ([26e568a](26e568a)) * **bigquery:** add zero-copy queryArrow API for Arrow VectorSchemaRoot streaming ([#14402](#14402)) ([b44ffe8](b44ffe8)) * **bigquery:** make BigQuery AutoCloseable with default no-op close method ([#14434](#14434)) ([00bf3de](00bf3de)) * **firestore:** add support for BSON types ([#13189](#13189)) ([8a123d9](8a123d9)) * **gax:** add ApiCallContext and request-level settings overloads to ResumableUploadCallable ([#14251](#14251)) ([e8cbd42](e8cbd42)) * **gax:** add globalTimeout settings field to ResumableUploadCallSettings ([#14253](#14253)) ([438cda6](438cda6)) * **gax:** add resumable upload error classification and retry algorithm ([#14419](#14419)) ([b70396d](b70396d)) * **gax:** add ResumableUploadCallable creation to Callables and HttpJsonCallableFactory ([#14242](#14242)) ([7de24de](7de24de)) * **gax:** implement baseline Callable and Future for resumable uploads ([#14241](#14241)) ([5a54db9](5a54db9)) * **generator:** add model flag and allowlist parser for resumable upload RPCs ([#14317](#14317)) ([acc1856](acc1856)) * **generator:** emit resumable upload client surface ([#14319](#14319)) ([a9fed00](a9fed00)) * **generator:** emit resumable upload settings and HttpJson upload stub ([#14321](#14321)) ([c122474](c122474)) * **generator:** enable resumable upload generation for showcase ([#14325](#14325)) ([f9ebd79](f9ebd79)) * **generator:** switch resumable upload specialized stubs to package private ([#14471](#14471)) ([0d4e875](0d4e875)) * **generator:** wire transport stub delegation to resumable upload stubs ([#14322](#14322)) ([cc4b980](cc4b980)) * **google/cloud/backupdr/v1beta:** add backupdr ([#14410](#14410)) ([a4a47da](a4a47da)) * **google/cloud/networkservices/v1beta1:** add networkservices ([#14407](#14407)) ([21c4955](21c4955)) * **pubsub:** add publish telemetry headers for publish attempt observability ([#14338](#14338)) ([c167ab8](c167ab8)) * **pubsub:** implement publish hedging to reduce tail latency ([#13735](#13735)) ([b302615](b302615)) * **spanner:** Support dynamic TLS certificate and key rotation for Spanner Omni ([#14456](#14456)) ([ffc745c](ffc745c)) * **storage/control:** add delete folder recursive sample ([#13642](#13642)) ([f4b1b46](f4b1b46)) * **storage/control:** add delete folder recursive sample ([#14397](#14397)) ([2c01d55](2c01d55)) ### Bug Fixes * **auth:** restore transportFactory upon deserialization in InternalAwsSecurityCredentialsSupplier ([#14340](#14340)) ([beea42f](beea42f)) * **bigquery-jdbc:** ensure row ordering in PCNT IT ([#14330](#14330)) ([a16f048](a16f048)) * **bigquery-jdbc:** fix htapi fallback due to permission logic ([#14418](#14418)) ([21e6dc8](21e6dc8)) * **bigquery-jdbc:** fix Timestamp assertions ([#14290](#14290)) ([533ba14](533ba14)) * **bigquery-jdbc:** handle null parameters in Storage Write API bulk inserts ([#14270](#14270)) ([dd2c41a](dd2c41a)), refs [#14066](#14066) * **bigquery-jdbc:** handle SQL NULLs in ResultSet primitive getters ([#14383](#14383)) ([8e464fe](8e464fe)), refs [#14371](#14371) * **bigquery:** default Arrow pagination stream location to US instead of global ([#14458](#14458)) ([2775eb1](2775eb1)) * **bigquery:** preserve page token and paginate correctly in Arrow query when maxResults is set ([#14469](#14469)) ([f5601f4](f5601f4)) * **bigquery:** use first page row count for Arrow query pagination offset ([#14466](#14466)) ([9d10dd0](9d10dd0)) * **bigtable:** don't notify config listeners while holding the manager lock ([#14294](#14294)) ([4426ccd](4426ccd)) * **bigtable:** fall back to classic path when per-RPC CallCredentials are set on session path ([#14477](#14477)) ([57bacb0](57bacb0)) * **bigtable:** fix abnormal session closures and scale-up in session pool ([#14431](#14431)) ([6361ecd](6361ecd)) * **biqguery:** fix undeclared QueryParameter wiring in QueryStatistics ([#14401](#14401)) ([64cf1d3](64cf1d3)) * **bom:** restore google-cloud-spanner-jdbc to libraries-bom ([#14362](#14362)) ([bc7be5e](bc7be5e)), refs [#14347](#14347) * **spanner:** honor maxAttempts and totalTimeout in streaming resume loop ([#14370](#14370)) ([305f47d](305f47d)) * **spanner:** only set snapshot isolation read timestamp for SI or optimistic txns in CloudClientExecutor ([#14346](#14346)) ([54c0d0f](54c0d0f)) * **spanner:** prevent statement cancellation race in AbstractBaseUnitOfWork ([#14283](#14283)) ([d9a8eef](d9a8eef)) * **spanner:** re-enable ITInstanceAdminTest on cloud-devel and cloud-staging ([#14281](#14281)) ([89a8268](89a8268)) ### Performance Improvements * **spanner:** stop re-parsing the request id on every RPC ([#14353](#14353)) ([46108f4](46108f4)) ### Documentation * Add a Http/Json Post-Quantum Cryptography Guide ([#13963](#13963)) ([fcc65b0](fcc65b0)) * **bigquery:** add QueryArrow code sample and document JDK 17+ JVM requirements ([#14437](#14437)) ([bd363f6](bd363f6)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: release-please[bot] <55107282+release-please[bot]@users.noreply.github.com>
Adds the ArrowDeserializer class which handles decoding serialized Arrow schemas and record batches into standard FieldValueList rows.