Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Update table metadata - #139

Merged
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr
Dec 4, 2023
Merged

Update table metadata #139
Fokko merged 24 commits into
apache:mainfrom
HonahX:table_metadata_update_pr

Conversation

@HonahX

@HonahXHonahX commented Nov 12, 2023

Copy link
Copy Markdown
Contributor

Fixes: #22

This PR:

  • Implements update_table_metadata which applies a list of TableUpdates on the current metadata
  • Adds unit tests in test_init.py

In the first version, update_table_metadata will support the following types of table updates:

This PR is one of the parts required to support_commit_table of catalogs except rest. A WIP PR containing a complete implementation of glue table commit can be found here: #140

@HonahXHonahX changed the title Implement table metadata update and table requirement validationImplement table metadata updateNov 12, 2023
@HonahXHonahX mentioned this pull request Nov 12, 2023
@HonahXHonahX changed the title Implement table metadata updateTable metadata updateNov 12, 2023
@HonahXHonahX changed the title Table metadata updateUpdate table metadata Nov 12, 2023
@Fokko

Copy link
Copy Markdown
Contributor

@HonahX This is high on my list. I'm OOO the rest of the week, I'll review this early next week since this is quite an important PR that needs some focus.

@Fokko
Fokko self-requested a review November 15, 2023 22:57

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good @HonahX, left some comments.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

class TableMetadataUpdateContext:
updates: List[TableUpdate]
last_added_schema_id: Optional[int]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't use this one?

Suggested change
last_added_schema_id: Optional[int]

@HonahXHonahXNov 22, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one case that we need this:
For SetCurrentSchemaUpdate, if the given schema_id is -1, we will set the current schema to last added schema by the stored last_added_schema_id.

I think later when we add support to SetDefaultSpecUpdate and SetDefaultSortOrderUpdate we can add last_added_spec_id and last_added_sort_order_id here too. WDYT?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to keep them separate. I think we might need to have some additional checks here such as, what happens if you add a column, and then revoke the column again. It will first create a new schema, with a new ID, and then it will reuse the old schema again.

  • 1: Schema(a: int), current_schema_id=1
    Add column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=2
    Drop column b:
  • 1: Schema(a: int), 2: Schema(a: int, b: int), current_schema_id=1

@HonahXHonahXDec 3, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for sharing the example! I'd like to clarify some points to ensure I fully understand your perspective. According to Java ref, It appears that lastAddedSchemaId refers to the most recently added schema in the current transaction.

The purpose of lastAddedSchemaId in Java is to support multiple AddSchema updates within a single transaction, allowing for efficient tracking of the latest schema addition. This is particularly beneficial if a schema added later in the transaction is the same as one added earlier. However, in Python, where we limit to one AddSchema per transaction, lastAddedSchemaId should either be the ID of the added schema or None, specifically when the added schema is identical to existing ones in the metadata.

The example you provided brings up an important consideration for our Python implementation: if the result schema of UpdateSchema is identical to a previously added schema in base_metadata, we should not add this new schema to the metadata. Instead, we only need to set the current schema ID to that of the previously existing one.
In Java, metadata builder will suppress the AddSchema update if the new schema is a duplicate of any previous schemas. Inspired by this, I think we can implement a similar mechanism in UpdateSchema.commit(). If the new schema is identical to an existing one, only a single setCurrentSchema should be added.

In conclusion, to enhance our approach to schema updates, I propose the following:

  1. In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.
  2. Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

What do you think – do the above reasoning and proposed changes seem good to you?

Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough explanation.

In UpdateSchema.commit(): We should check if new_schema is identical to any existing schemas before adding an AddSchema and a SetCurrentSchema. If they are identical, only a SetCurrentSchema(existing_schema_id) is necessary.

Yes, this makes a lot of sense to me.

Regarding lastAddedSchemaId: It’s crucial to ensure that if it's not None, it points to a schema added earlier in the current transaction. Since Python transactions contain only one AddSchema, this should be straightforward. We could also consider adding a uniqueness check in update_table_metadata to confirm that the updates are from a single transaction. Also, instead of having a lastAddedSchemaId, we may add a method called last_added_schema() to extract the last added schema from the update context.

This also makes sense to me. However one thing to note. The lastAddedSchemaId sounds to me like an implementation detail from Java. In PyIceberg the current situation is simpler as you explained, so we could just do max(schema.id from tbl.schemas) + 1.

You sparked my interest here, let me double check this on the REST side as well 👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for raising this, this is currently not handled correctly on the rest side either.

I've created a PR here as a suggestion, but also feel free to fix it in your PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX can you rebase? I think this one is ready to go!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Fokko Thanks for reviewing! I've pushed my updates (mostly on SetCurrentSchemaUpdate) and rebased the PR.

Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py
Comment on lines +408 to +409
updated_metadata_data = copy(base_metadata.model_dump())
updated_metadata_data["format-version"] = update.format_version

@FokkoFokkoNov 20, 2023

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While this is a very safe way of doing the copy, it is also rather expensive since we convert everything to a Python dict, and then create a new object again. Pydantic has the model_copy argument that seems to do what we're looking for:

Suggested change
updated_metadata_data=copy(base_metadata.model_dump())
updated_metadata_data["format-version"] =update.format_version
updated_metadata_data=base_metadata.model_copy(**{"format-version": update.format_version})

We could construct a dict where we add all the changes (for the more complicated updated below), and then call model_copy for each update.

This will make a shallow copy by default (which I think is okay, since the model is immutable).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the explanation! I am a little worried about the immutability of table metadata. I think Pydantic's frozen config does not prevent updates to list, dict, etc. If we make shallow copy of list fields in metadata and later some code mistakenly alter the list (e.g. append something) in the updated metadata, the effect will be populated to the base_metadata too and it may be hard to detect. Do you think this might be a problem?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case, we can be cautious and set deep=true. I would love to see some tests that validate the behavior. Those should be easy to add.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion! I've just implemented the suggested change on my end, but I'm still in the process of building the tests for shallow vs deep copy. Given that the current PR already contains lots of change, do you think it might be a good idea to make the model_copy transfer in a separate, follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @HonahX. I've created a new issue here: #179

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py
Comment threadpyiceberg/table/__init__.py Outdated
Comment threadpyiceberg/table/__init__.py Outdated

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@HonahX Almost there, looking good 👍

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @HonahX Thanks for picking this up!

@Fokko
Fokko merged commit 8330610 into apache:mainDec 4, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ability to the write Metadata JSON

2 participants

@HonahX@Fokko