Skip to content

samples(bigquery-storage): add Arrow query results samples for query_and_wait and read_rows - #18126

Open
alextolpin wants to merge 2 commits into
googleapis:mainfrom
alextolpin:arrow_samples
Open

samples(bigquery-storage): add Arrow query results samples for query_and_wait and read_rows#18126
alextolpin wants to merge 2 commits into
googleapis:mainfrom
alextolpin:arrow_samples

Conversation

@alextolpin

Copy link
Copy Markdown
Contributor

Description

Adds documentation code snippets and system tests demonstrating high-performance query result retrieval in Apache Arrow format with LZ4 compression using the BigQuery Storage API:

  1. query_and_wait() with Arrow format & LZ4 frame compression (query_and_wait_arrow.py):

    • Demonstrates calling client.query_and_wait() with query_results_format=enums.QueryResultsFormat.ARROW and compression_codec=enums.QueryResultsCompressionCodec.LZ4_FRAME.
    • Returns results as an iterable of pyarrow.RecordBatch via results.to_arrow_iterable().
    • Wrapped in region tag: [START bigquerystorage_query_and_wait_arrow].
  2. Direct read_rows on query job default stream (read_rows_query_job.py):

    • Demonstrates directly reading query results using BigQueryReadClient.read_rows against the job stream projects/{project}/locations/{location}/jobs/{job_id}/streams/_default.
    • Deserializes schema and record batches safely via pyarrow.ipc.
    • Wrapped in region tag: [START bigquerystorage_read_rows_query_job].
  3. Tests & Dependencies:

    • Added query_and_wait_arrow_test.py and read_rows_query_job_test.py verifying batch iteration and schema types.
    • Added pyarrow dependency pins to samples/snippets/requirements.txt.

Follow-up to #18027
Related to #18047

Checklist

  • Ensure the tests and linter pass
  • Code coverage does not decrease (if any source code was changed)
  • Appropriate docs were updated (if necessary)

@alextolpin
alextolpin requested review from a team as code ownersAugust 16, 2026 17:59
@alextolpin
alextolpin requested review from sindhuvy and removed request for a teamAugust 16, 2026 17:59
@snippet-bot

Copy link
Copy Markdown

Here is the summary of changes.

You are about to add 2 region tags.

This comment is generated by snippet-bot.
If you find problems with this result, please file an issue at:
https://github.com/googleapis/repo-automation-bots/issues.
To update this comment, add snippet-bot:force-run label or use the checkbox below:

  • Refresh this comment

@gemini-code-assistgemini-code-assistBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces two new code snippets and their corresponding tests to demonstrate reading BigQuery query results in Apache Arrow format. Specifically, query_and_wait_arrow.py uses the query_and_wait method to fetch Arrow RecordBatches directly, while read_rows_query_job.py initiates a query job and streams the results using the BigQuery Storage Read API. A critical issue was identified in read_rows_query_job.py where the query job is started asynchronously, but the code immediately attempts to read from the stream without waiting for the job to complete. It is recommended to call job.result() to ensure the query finishes before reading.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@alextolpin