Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Arrow: Allow missing field-ids from Schema - #183

Closed
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional
Closed

Arrow: Allow missing field-ids from Schema#183
Fokko wants to merge 4 commits into
apache:mainfrom
Fokko:fd-make-field-ids-optional

Conversation

@Fokko

@FokkoFokko commented Dec 5, 2023

Copy link
Copy Markdown
Contributor

No description provided.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from e5923dd to 405d36cCompareDecember 5, 2023 12:43

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To clarify my understanding, I think this fallback is expected to function correctly only if the table schema's field IDs are generated in the same way as in the visitor or assign_fresh_field_id. This should generally hold true, as the official Iceberg API always reassigns schema field IDs when creating a table. Please let me know if I've misunderstood anything here. Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
missing_is_metadata: Optional[bool]

def __init__(self) -> None:
self.counter = count()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we use count(1) here since we start from 1 inassign_fresh_schema_ids

class_SetFreshIDs(PreOrderSchemaVisitor[IcebergType]):
"""Traverses the schema and assigns monotonically increasing ids."""
reserved_ids: Dict[int, int]
def__init__(self, next_id_func: Optional[Callable[[], int]] =None) ->None:
self.reserved_ids= {}
counter=itertools.count(1)
self.next_id_func=next_id_funcifnext_id_funcisnotNoneelselambda: next(counter)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the limited context here, it will skip the fields if it doesn't have an ID:
image

Which is kind of awkward.

@FokkoFokkoDec 7, 2023

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using assign_fresh_schema_ids, since that one is a pre-order, and the default is post-order. I've updated the code, let me know what you think! Appreciate the review!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data. This is really great! Out of curiosity, are there any additional use-cases where this PR might be beneficial?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I happen to have an iceberg table (migrated from delta lake) whose parquet files contain no field-id. With this change, I am now able to use pyiceberg to read its data.

There is already a way to assign field IDs when they are not in a data file, using a name mapping. All reads that need to infer field IDs must use a name mapping rather than assigning IDs per data file.

@Fokko
Fokkoforce-pushed the fd-make-field-ids-optional branch from d507fcd to 27017cfCompareDecember 7, 2023 15:20

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my current understanding, this Pull Request can be a temporal solution for managing parquet files lacking field-ids, until the Name-mapping feature is implemented. Overall, it seems good to me! Just have one suggestion on adding more details in the warning message

Additionally, I want to point out some prerequisites for the field-id assignment process to work when reading tables:

  1. The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.
  2. The column order in the parquet file should align with that in the table_schema.

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

Thanks!

Comment threadpyiceberg/io/pyarrow.py Outdated
@Fokko

Fokko commented Dec 9, 2023

Copy link
Copy Markdown
ContributorAuthor

Thanks for the Java context here! Appreciate it!

The table_schema must be processed with assign_field_id (or its equivalent function in Java) before being written to the table.

This will still be the case, but this aims for when we're just reading a Parquet file (that doesn't have field-id metadata), and turning that dataframe into a table. So this is more on the read than the write side. Currently, if you transform a field, it will just drop the fields which is not what we want.

The column order in the parquet file should align with that in the table_schema.

I agree, we could see if we can have a way to wire up the names. That's a great idea! Could you create an issue on that?

If these pre-reqs are not met, we might encounter errors during column binding and data reading. I think this is the reason that we want Name-mapping finally. Please let me know if there's any aspect of this I might be misunderstanding.

I don't see the link with name-mapping. Name mapping will provide an external mapping of column names to ID. But this is more about writing new data to a new Iceberg table where there are no IDs.

Just thinking out loud from a practical perspective. If you retrain a model, you have results, and you want to overwrite an existing table that has the same column, then you want to wire up those names then we might want to defer the assignment of IDs until we write. I think we can elaborate on this.

@HonahX

Copy link
Copy Markdown
Contributor

Thanks for the detailed explanation! I was originally looking at how this PR helps with reading, especially focusing on the code here:

ifmetadata:=physical_schema.metadata:
schema_raw=metadata.get(ICEBERG_SCHEMA)
# TODO: if field_ids are not present, Name Mapping should be implemented to look them up in the table schema,
# see https://github.com/apache/iceberg/issues/7451
file_schema=Schema.model_validate_json(schema_raw) ifschema_rawisnotNoneelsepyarrow_to_schema(physical_schema)
. That's why I brought up name-mapping.

But now I totally get the write-side importance too. Based on your suggestion, I've opened Issue #199. Would love your thoughts on it or any suggestions you might have!

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@sungwysungwy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM as well. I was running into an issue reading iceberg tables that had data files side-loaded into the Iceberg Table with add_files migration procedure, and this PR will resolve the issue.

/lib/python3.10/site-packages/pyiceberg/io/pyarrow.py in _get_field_id(field)
703 def _get_field_id(field: pa.Field) -> Optional[int]:
704 for pyarrow_field_id_key in PYARROW_FIELD_ID_KEYS:
--> 705 if field_id_str := field.metadata.get(pyarrow_field_id_key):
706 return int(field_id_str.decode())
707 return None
AttributeError: 'NoneType' object has no attribute 'get'

@rdblue

Copy link
Copy Markdown
Contributor

Per my comment, I'm -1 if the intent of this PR is to read data files without IDs. That must use a name mapping in order to be safe!

@FokkoFokko added this to the PyIceberg 0.6.0 release milestone Dec 13, 2023
@sungwysungwy mentioned this pull request Dec 17, 2023
3 tasks
@FokkoFokko closed this Dec 19, 2023
@HonahXHonahX mentioned this pull request Dec 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Fokko@HonahX@rdblue@sungwy