Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Python: Fine-tune the API - #5672

Merged
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs
Sep 20, 2022
Merged

Python: Fine-tune the API#5672
rdblue merged 10 commits into
apache:masterfrom
Fokko:fd-update-python-docs

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

The API wasn't consistent everywhere. Now the ids will just initialize at 1, so the user doesn't have to do this.

@Fokko
Fokko marked this pull request as draft August 30, 2022 19:58
@Fokko

Fokko commented Aug 30, 2022

Copy link
Copy Markdown
ContributorAuthor

Waiting for #5627

The API wasn't consistent everywhere. Now the ids will just initialize
at 1, so the user doesn't have to do this.
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 75d7e3a to dd28be8CompareSeptember 1, 2022 20:02
@Fokko
Fokko marked this pull request as ready for review September 1, 2022 20:03
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 0f32ba3 to a0806cdCompareSeptember 2, 2022 16:36
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-api-intro.md Outdated
Comment threaddocs/python-feature-support.md
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threaddocs/python-quickstart.md Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
"""A list of schemas, stored as objects with schema-id."""

current_schema_id: int = Field(alias="current-schema-id", default=DEFAULT_SCHEMA_ID)
current_schema_id: int = Field(alias="current-schema-id", default=INITIAL_SCHEMA_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These probably shouldn't have defaults because they need to be explicitly set to some ID that exists in the list of schemas, specs, or sort orders.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that defaulting the ID when creating a new schema, spec, or order is fine. But I don't think it is a good idea to default it here. At this point, we no longer have users constructing metadata by hand and we want to make sure that we're setting the ID correctly. If we re-create a schema for a new table metadata object, then we should also set the current schema ID to that schema's ID rather than relying on the same default in two places. That way if we ever change the default assignment we don't break tables.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point. I've removed this set explicitly when we get a v1 metadata.

Comment threadpython/pyiceberg/table/partitioning.py Outdated

spec_id: int = Field(alias="spec-id")
fields: Tuple[PartitionField, ...] = Field(default_factory=tuple)
spec_id: int = Field(alias="spec-id", default=INITIAL_PARTITION_SPEC_ID)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I'd prefer to handle ID assignment manually rather than defaulting. Defaulting seems to bring in complexity because if we forget to pass along an ID somewhere, it would cause problems.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel that we don't should really expose this to the user. For example, when we create a new table, we re-assign the IDs anyway (using the assign fresh IDs logic).
If we follow the Java API, and we have something similar to updateSpec: https://github.com/apache/iceberg/blob/master/api/src/main/java/org/apache/iceberg/Table.java#L165-L171 Then we can just take the next ID. What do you think of this?

@rdbluerdblueSep 5, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds reasonable to me. I think we just need to make sure that reassignment is correct!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will definitely involve a lot of testing 👍🏻

Comment threadpython/pyiceberg/table/partitioning.py Outdated
Comment threadpython/pyiceberg/table/sorting.py
Comment threadpython/pyproject.toml Outdated
Comment threaddocs/python-api-intro.md Outdated
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 61147bd to 8fa8d88CompareSeptember 5, 2022 20:29
@FokkoFokko changed the title Python: Update docs and fine-tune the APIPython: Fine-tune the APISep 8, 2022
@Fokko

Fokko commented Sep 8, 2022

Copy link
Copy Markdown
ContributorAuthor

Split out the changes to the docs to #5727

@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 446d94e to 6f49ae3CompareSeptember 19, 2022 09:23
@Fokko
Fokkoforce-pushed the fd-update-python-docs branch from 6f49ae3 to 0614692CompareSeptember 19, 2022 13:58
@Fokko

Copy link
Copy Markdown
ContributorAuthor

@rdblue I've resolved the merge conflicts, would you have time for another pass? Thanks!


Args:
order_id (int): The id of the sort-order. To keep track of historical sorting
order_id (int): An unique id of the sort-orderof a table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need "of a table" -- that assumes the context that uses the sort order.

location=None,
partition_spec=PartitionSpec(
spec_id=1, fields=(PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"),)
PartitionField(source_id=1, field_id=1000, transform=TruncateTransform(width=3), name="id"), spec_id=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like the mock causes the result to not match the request. We should start testing against the REST catalog servlet as soon as we can.

@rdblue
rdblue merged commit b8a796e into apache:masterSep 20, 2022
@rdblue

Copy link
Copy Markdown
Contributor

Looks good. There were a couple minor things, but those aren't blockers.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Fokko@rdblue