WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

WIP: [AIRFLOW-1894] Google cloud bigquery - #4607

Closed
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery
Closed

WIP: [AIRFLOW-1894] Google cloud bigquery#4607
jmcarp wants to merge 1 commit into
apache:masterfrom
jmcarp:google-cloud-bigquery

Conversation

@jmcarp

@jmcarpjmcarp commented Jan 28, 2019

Copy link
Copy Markdown
Contributor

Make sure you have checked all steps below.

Jira

  • My PR addresses the following Airflow Jira issues and references them in the PR title. For example, "[AIRFLOW-XXX] My Airflow PR"

Description

  • Here are some details about my PR, including screenshots of any UI changes:

Tests

  • My PR adds the following unit tests OR does not need testing for this extremely good reason:

Commits

  • My commits all reference Jira issues in their subject lines, and I have squashed multiple commits if they address the same issue. In addition, my commits follow the guidelines from "How to write a good git commit message":
    1. Subject is separated from body by a blank line
    2. Subject is limited to 50 characters (not including Jira issue reference)
    3. Subject does not end with a period
    4. Subject uses the imperative mood ("add", not "adding")
    5. Body wraps at 72 characters
    6. Body explains "what" and "why", not "how"

Documentation

  • In case of new functionality, my PR adds documentation that describes how to use it.
    • When adding new operators/hooks/sensors, the autoclass documentation generation needs to be added.
    • All the public functions and the classes in the PR contain docstrings that explain what it does

Code Quality

  • Passes flake8

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 4 times, most recently from 94a6905 to d131b00CompareJanuary 30, 2019 04:25
@jmcarpjmcarp changed the title [AIRFLOW-3776] Google cloud bigquery[AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 8bd9564 to dc502cbCompareJanuary 30, 2019 05:03
@kaxilkaxil changed the title [AIRFLOW-1894] Google cloud bigqueryWIP: [AIRFLOW-1894] Google cloud bigqueryJan 30, 2019
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 2 times, most recently from 89b048d to 292d467CompareJanuary 31, 2019 01:43
@codecov-io

codecov-io commented Jan 31, 2019

Copy link
Copy Markdown

Codecov Report

Merging #4607 into master will increase coverage by 0.23%.
The diff coverage is 47.05%.

Impacted file tree graph

@@ Coverage Diff @@## master #4607 +/- ##
==========================================
+ Coverage 74.3% 74.53% +0.23% 
==========================================
Files 426 426 Lines 27867 27577 -290 ==========================================
- Hits 20706 20554 -152 + Misses 7161 7023 -138
Impacted FilesCoverage Δ
...rflow/contrib/operators/bigquery_check_operator.py0% <ø> (ø)⬆️
airflow/contrib/operators/gcs_to_bq.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_get_data.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_bigquery.py0% <0%> (ø)⬆️
...ontrib/operators/bigquery_table_delete_operator.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_to_gcs.py0% <0%> (ø)⬆️
airflow/contrib/operators/bigquery_operator.py95.93% <100%> (+2.36%)⬆️
airflow/contrib/hooks/bigquery_hook.py60.69% <47.29%> (+2.62%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update f07f3a8...d903f24. Read the comment docs.

@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch 5 times, most recently from 7562a62 to 8698ae7CompareFebruary 2, 2019 02:13
@jmcarp
jmcarpforce-pushed the google-cloud-bigquery branch from 8698ae7 to d903f24CompareFebruary 2, 2019 02:18
@jmcarp

Copy link
Copy Markdown
ContributorAuthor

@kaxil@potiuk: tests are passing, so this is ready for review when you have time. A few differences to point out:

  • The current python bigquery client library implements the dbapi connection and cursor interfaces, so we can drop our custom implementations
  • The current client also allows for polling a job by calling its result method, so we can also drop our custom polling logic
  • Because the current client includes classes for most job options that include basic validation, we're also able to drop some custom validation
  • Extra methods on the bigquery cursor have been moved to the hook

@jmcarp

Copy link
Copy Markdown
ContributorAuthor

Ping @Fokko@feng-tao@mik-laj (not sure who knows this code best). I'm hoping to add some features after this is ready, so would be great to get feedback when you all have time.

sql,
bigquery_conn_id='bigquery_default',
use_legacy_sql=True,
use_legacy_sql=False,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't change this behaviour

allow_jagged_rows=self.allow_jagged_rows,
src_fmt_configs=self.src_fmt_configs,
labels=self.labels
external_config_options=self.external_config_options,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will have to be for 2.0 as it will break things. Also, it needs to be backward-compatible to make updation smooth from 1.X to 2.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Haven't yet reviewed the entire PR but schemed through it. Will try to find some time to look at it more thoroughly

@Fokko

Copy link
Copy Markdown
Contributor

I'm seeing some good things in this PR, and it looks like it will simplify the logic. Any plans of moving this forward @jmcarp ?

:type project_id: str
"""
service = self.get_service()
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you could use fallback_to_default_project_id decorator. It has additional logic to raise the exception if none of the project_id s is specified. and you could remove this if altogether then. It forces to use keyword parameters though.

project_id=project,
use_legacy_sql=self.use_legacy_sql,
def get_client(self, project_id=None):
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as below - fallback_to_default_project_id decorator is nicer way I think

}
})
if external_config_options is not None:
if not isinstance(external_config_options, type(external_config.options)):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should use dict explicitly here. What if external_config.options are None (seem to be default).

https://cloud.google.com/bigquery/docs/locations#specifying_your_location
:type location: str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again - I think using fallback decorator is nicer as it keeps the project_id logic in one place.

passed to BigQuery
:type labels: dict
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here with decorator.

time_partitioning. The order of columns given determines the sort order.
:type cluster_fields: list of str
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here.

"""
# check to see if the table exists
table_id = table_resource['tableReference']['tableId']
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And here :)

even if any insertion errors occur.
:type fail_on_error: bool
"""
project_id = project_id if project_id is not None else self.project_id

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here too :)

private_key=private_key)

def table_exists(self, project_id, dataset_id, table_id):
def table_exists(self, dataset_id, table_id, project_id=None):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi, thanks for putting this together, this is great stuff!

Regarding these convenience methods, I'd argue that it would be better if the functionality would be implemented on the connection object returned by get_conn, and the here we just delegate, like this:

deftable_exists(self, dataset_id, table_id, project_id=None):
returnself.get_conn().table_exists(dataset_id, table_id, project_id)

The reason is that sometimes it is very useful to have explicit control over when and where connections get created, for example when using multi-threading for I/O optimization. We've also had code that spuriously crashed because it was creating too many new connections (implicitly in such convenience methods), and at some point the authentication requests hit a rate-limit.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp -> are you still working on this? As I am now committer, I am happy to review that one as well as soon as it gets rebased (if it's still something you want to add).

@kaxil

Copy link
Copy Markdown
Member

@jmcarp Are you still working on this?

@mik-lajmik-laj added the provider:google Google (including GCP) related issues label Jul 22, 2019
@mik-laj

Copy link
Copy Markdown
Member

@jmcarp Can I help you with this? This looks very interesting to me.

@mik-laj

Copy link
Copy Markdown
Member

Hi.

I made a change in the base class - GoogleCloudBaseHook. Your PR may need to be changed. Could you do rebase?

Cheers

Refenence:
#5907

@potiuk

Copy link
Copy Markdown
Member

@jmcarp - is it possible to rebase/complete the work on this one ? It blocks us from moving operators/hook from contrib to core - after we move the operators/hooks it will be much more difficult to merge.

@potiuk

Copy link
Copy Markdown
Member

Hey @jmcarp - are you still doing it? Or should we close that one?

@stale

staleBot commented Nov 20, 2019

Copy link
Copy Markdown

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stalestaleBot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Nov 20, 2019
@stalestaleBot closed this Nov 27, 2019
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

provider:googleGoogle (including GCP) related issuesstaleStale PRs per the .github/workflows/stale.yml policy file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@jmcarp@codecov-io@Fokko@potiuk@kaxil@mik-laj@bjoernpollex-sc