Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Improve the InMemory Catalog Implementation - #289

Merged
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test
Mar 13, 2024
Merged

Improve the InMemory Catalog Implementation #289
Fokko merged 24 commits into
apache:mainfrom
kevinjqliu:kevinliu/test

Conversation

@kevinjqliu

@kevinjqliukevinjqliu commented Jan 20, 2024

Copy link
Copy Markdown
Contributor

Issue #293

Improve the InMemory Catalog implementation.

In this PR:

  • Implement In-Memory catalog’s create_table and _commit_table function
  • Added default warehouse location which can be on the local file system (defaults to /tmp/warehouse)
  • Change test_base.pytest_console.py to write to a temporary file location on the local file system using tmp_path from pytest
  • Fix test_commit_table from tests/catalog/test_base.py, issue described in schema_id not incremented during schema evolution  #290

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 3 times, most recently from 9242e69 to c09e7e9CompareJanuary 22, 2024 00:32
@kevinjqliukevinjqliu changed the title Kevinliu/testInMemory Catalog Implementation Jan 22, 2024
Comment threadpyiceberg/catalog/in_memory.py Outdated
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER, SortOrder
from pyiceberg.typedef import EMPTY_DICT

DEFAULT_WAREHOUSE_LOCATION = "file:///tmp/warehouse"

@kevinjqliukevinjqliuJan 22, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by default, write on disk to /tmp/warehouse

Comment threadpyiceberg/catalog/in_memory.py Outdated
super().__init__(name, **properties)
self.__tables = {}
self.__namespaces = {}
self._warehouse_location = properties.get(WAREHOUSE, None) or DEFAULT_WAREHOUSE_LOCATION

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can pass a warehouse location using properties. warehouse location can be another fs such as s3

@pytest.fixture
def catalog() -> InMemoryCatalog:
return InMemoryCatalog("test.in.memory.catalog", **{"test.key": "test.value"})
def catalog(tmp_path: PosixPath) -> InMemoryCatalog:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added ability to write to temporary files for testing, which is then automatically cleaned up

@kevinjqliu
kevinjqliuforce-pushed the kevinliu/test branch 2 times, most recently from 4842d91 to a56838dCompareJanuary 22, 2024 05:06
Comment threadtests/catalog/test_base.py Outdated
assert response.metadata.table_uuid == given_table.metadata.table_uuid
assert len(response.metadata.schemas) == 1
assert response.metadata.schemas[0] == new_schema
assert given_table.metadata.current_schema_id == 1

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kevinjqliu
kevinjqliu marked this pull request as ready for review January 22, 2024 05:15

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your contribution! @kevinjqliu I left some initial comments below.

Comment threadpyiceberg/table/__init__.py Outdated
Comment threadtests/catalog/test_base.py
Comment threadpyiceberg/catalog/in_memory.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work @kevinjqliu

Should we also add this catalog to the tests in tests/integration/test_reads.py?

Comment threadpyiceberg/catalog/__init__.py Outdated
CatalogType.GLUE: load_glue,
CatalogType.DYNAMODB: load_dynamodb,
CatalogType.SQL: load_sql,
CatalogType.MEMORY: load_memory,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you also add this one to the docs: https://py.iceberg.apache.org/configuration/ With a warning that this is just for testing purposes only.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if identifier in self.__tables:
raise TableAlreadyExistsError(f"Table already exists: {identifier}")
else:
if namespace not in self.__namespaces:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other implementations don't auto-create namespaces, however I think it is fine for the InMemory one.

Comment threadpyiceberg/catalog/in_memory.py Outdated
if not location:
location = f'{self._warehouse_location}/{"/".join(identifier)}'

metadata_location = f'{self._warehouse_location}/{"/".join(identifier)}/metadata/metadata.json'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like we don't write the metadata here, but we write it below at the _commit method

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, the actual writing is done by _commit_table below, but the path of the metadata location is determined here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, but I'm a bit confused here. If I just want to create the table without inserting any data:

catalog.create_table(schema, ....)

I still expect a new metadata.json file to be found at the table location without any call to _commit_table. But that does not seem to be created by the InMemory catalog now. Is there a reason that we choose this behavior?

In the previous implementation no file is written. But since we have updated _commit_table to write the metadata file, I think it more reasonable to make create_table aligned with other production implementation. WDYT?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gotcha, that makes sense!

Comment threadtests/catalog/test_base.py

@HonahXHonahX left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM! Thanks for updating this to a formal implementation and adding the doc.
I just have one more comment about create_table


def text(self, response: str) -> None:
Console().print(response)
Console(soft_wrap=True).print(response)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

some test_console.py outputs are too long and end up with an extra \n in the middle of the string, causing tests to fail

identifier=TEST_TABLE_IDENTIFIER,
schema=pyarrow_schema_simple_without_ids,
location=TEST_TABLE_LOCATION,
partition_spec=TEST_TABLE_PARTITION_SPEC,

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@syun64 FYI, I realized that the TEST_TABLE_PARTITION_SPEC here breaks this test.

TEST_TABLE_PARTITION_SPEC = PartitionSpec(PartitionField(name="x", transform=IdentityTransform(), source_id=1, field_id=1000))

The partition field's source_id here is 1, but in create_table the schema's field_ids are all -1 due to _convert_schema_if_needed

So assign_fresh_partition_spec_ids fails

original_column_name=old_schema.find_column_name(field.source_id)
iforiginal_column_nameisNone:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @kevinjqliu thank you for flagging this 😄 I think '-1' ID discrepancy is the symptom of the issue that makes the issue easy to understand, just as we decided in #305 (comment)

The root cause of the issue I think is that we are introducing a way for non-ID's schema (PyArrow Schema) to be used as an input into create_table, while not supporting the same for partition_spec and sort_order (PartitionField and SortField both require field IDs as inputs).

So I think we should update both assign_fresh_partition_spec_ids and assign_fresh_sort_order_ids to support field look up by name.

@Fokko - does that sound like a good way to resolve this issue?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created #338 to track this issue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you @syun64 that for creating tables having to look up the IDs is not ideal. Probably that API has to be extended at some point.

But for the metadata (and also how Iceberg internally tracks columns, since names can change; IDs not), we need to track it by ID. I'm in doubt if assigning -1 was the best idea because that will give you a table that you cannot work with. Thanks for creating the issue, and let's continue there.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good @Fokko 👍 and thanks again for flagging this @kevinjqliu !

@kevinjqliukevinjqliu mentioned this pull request Jan 31, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

hey @Fokko / @HonahX do you mind taking a look at this again?

@kevinjqliukevinjqliu changed the title InMemory Catalog Implementation Improve the InMemory Catalog Implementation Mar 1, 2024
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@Fokko As we discussed in #293, let's not create yet another catalog.
I moved the changes back to test_base.py where the In-Memory catalog was originally. This PR improves the implementation along with a bunch of testing improvements

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

Comment threadmkdocs/docs/configuration.md Outdated
Comment threadmkdocs/docs/configuration.md Outdated
Comment threadtests/catalog/test_base.py Outdated
kevinjqliuand others added 3 commits March 5, 2024 10:57
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
Co-authored-by: Fokko Driesprong <fokko@apache.org>
@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

Thanks for the suggestions, @Fokko

@kevinjqliu
kevinjqliu requested a review from FokkoMarch 9, 2024 16:50
@bitsondatadev

Copy link
Copy Markdown
Contributor

Thanks for working on this @kevinjqliu. The issues was created a long time ago, before we had the SqlCatalog with sqlite support. Sqlite can also work in memory rendering the InMemoryCatalog obsolete. Having two in-memory implementations in the codebase adds additional complexity in the codebase. My suggestion would be to replace the MemoryCatalog with the SqlCatalog. WDYT?

@kevinjqliu, this was likely answered offline and I suppose there was a reason to continue working here. I also am curious to know if this catalog still makes sense with inmem sqlite?

@kevinjqliu

Copy link
Copy Markdown
ContributorAuthor

@bitsondatadev

I also am curious to know if this catalog still makes sense with inmem sqlite?

We agreed to not move this implementation to production. See #289 (comment)
Instead, this PR is used to improve the InMemory catalog implementation in tests and use it to improve other tests

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for working on this @kevinjqliu, this is a great improvement 👍

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Mar 13, 2024
@Fokko
Fokko merged commit 36a505f into apache:mainMar 13, 2024
@kevinjqliu
kevinjqliu deleted the kevinliu/test branch March 13, 2024 15:01
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@kevinjqliu@Fokko@bitsondatadev@sungwy@HonahX