Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Python: Reassign schema/partition-spec/sort-order ids - #5627

Merged
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table
Aug 31, 2022
Merged

Python: Reassign schema/partition-spec/sort-order ids #5627
Fokko merged 7 commits into
apache:masterfrom
Fokko:fd-fresh-ids-when-creating-a-table

Conversation

@Fokko

Copy link
Copy Markdown
Contributor

When creating a new schema.

Also created a type alias called TableMetadata that replaces the Union[TableMetadataV1, TableMetadataV2] annotation.

Resolves#5468

Comment threadpython/pyiceberg/schema.py
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/schema.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
Comment threadpython/pyiceberg/table/metadata.py Outdated
def assign_fresh_partition_spec_ids(spec: PartitionSpec, schema: Schema) -> PartitionSpec:
partition_fields = []
for pos, field in enumerate(spec.fields):
schema_field = schema.find_field(field.name)

@rdbluerdblueAug 24, 2022

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the partition field name, not a schema field name. The schema field must be looked up by source_id. This method needs both the original schema and the fresh schema. The original schema is used to get field names and then the fresh schema is used to look up the new source ID.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great catch! 👍🏻

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko, looks like this hasn't been fixed yet, so I'm reopening the thread.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, that slipped through somehow

Comment threadpython/pyiceberg/table/sorting.py Outdated
@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from dc0dbc5 to 2df8ceeCompareAugust 25, 2022 08:09
Comment threadpython/pyiceberg/table/metadata.py
"""Visit a PrimitiveType"""


class PreOrderSchemaVisitor(Generic[T], ABC):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this pre-order? In Java we called it CustomOrder because you can choose when to visit children by accessing the callable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is pre-order traversal since we start at the root and then move to the leaves. In order is a bit less intuitive since it is not a binary tree. You could also do a reverse in-order, but not sure if we need that. We can also call it CustomOrder if you have a strong preference, but I think pre-order is the most logical way of using this visitor.

Comment threadpython/pyiceberg/schema.py Outdated
return next(self.counter)

def schema(self, schema: Schema, struct_result: Callable[[], StructType]) -> Schema:
return Schema(*struct_result().fields, identifier_field_ids=schema.identifier_field_ids)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this re-map the identifier field IDs since it is returning a new schema?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we do that in the function itself:

defassign_fresh_schema_ids(schema: Schema) ->Schema:
"""Traverses the schema, and sets new IDs"""schema_struct=pre_order_visit(schema.as_struct(), _SetFreshIDs())
fresh_identifier_field_ids= []
new_schema=Schema(*schema_struct.fields)
forfield_idinschema.identifier_field_ids:
original_field_name=schema.find_column_name(field_id)
iforiginal_field_nameisNone:
raiseValueError(f"Could not find field: {field_id}")
fresh_field=new_schema.find_field(original_field_name)
iffresh_fieldisNone:
raiseValueError(f"Could not lookup field in new schema: {original_field_name}")
fresh_identifier_field_ids.append(fresh_field.field_id)
returnnew_schema.copy(update={"identifier_field_ids": fresh_identifier_field_ids})

This is because we first want to know all the IDs

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I've refactored this because we need to build a map anyway 👍🏻

Comment threadpython/pyiceberg/schema.py Outdated

def field(self, field: NestedField, field_result: Callable[[], IcebergType]) -> IcebergType:
return NestedField(
field_id=self._get_and_increment(), name=field.name, field_type=field_result(), required=field.required, doc=field.doc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to visit children before visiting the next field. If you're trying to match the behavior of assignment in Java, you'd need to increment the counter for each field and then visit children.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missed that one, thanks! Just updated the code and tests

Comment threadpython/pyiceberg/schema.py
) -> TableMetadata:
fresh_schema = assign_fresh_schema_ids(schema)
fresh_partition_spec = assign_fresh_partition_spec_ids(partition_spec, fresh_schema)
fresh_sort_order = assign_fresh_sort_order_ids(sort_order, schema, fresh_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do the "fresh" methods always reset schema_id, spec_id, and order_id?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only when you create TableMetadata out of it (when creating a new table). And it resets if it isn't 1.

@Fokko
Fokkoforce-pushed the fd-fresh-ids-when-creating-a-table branch from ad95028 to bee1c81CompareAugust 30, 2022 19:45
@FokkoFokko mentioned this pull request Aug 30, 2022

@FokkoFokko left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll move this forward. Let me know if there is anything that you would like to see changed. There are two followups that I'd like to do:

  • Remove the pre-validators because they are confusing and error prone
  • Smooth out the API for the docs

@Fokko
Fokko merged commit 08bb3e2 into apache:masterAug 31, 2022
@Fokko
Fokko deleted the fd-fresh-ids-when-creating-a-table branch August 31, 2022 17:28
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: Re-assign IDs in when creating a table

2 participants

@Fokko@rdblue