Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Support for REPLACE TABLE operation - #433

Closed
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select
Closed

Support for REPLACE TABLE operation#433
anupam-saini wants to merge 31 commits into
apache:mainfrom
anupam-saini:as-replace-table-as-select

Conversation

@anupam-saini

@anupam-sainianupam-saini commented Feb 15, 2024

Copy link
Copy Markdown
Contributor

Closes#281
API proposal (from PR feedback):

table = catalog.create_or_replace_table(identifier, schema, location, partition_spec, sort_order, properties)

TODO:

  • Update schema
  • Update partition spec
  • Update sort order
  • Update location
  • Update table properties

Comment threadpyiceberg/schema.py Outdated
@Fokko

Copy link
Copy Markdown
Contributor

@anupam-saini Thanks for working on this. I'm not sure if the following API is where people would expect it:

withtable.transaction() astransaction:
transaction.replace_table_with(new_table)

Specially because this is an unsafe operation that breaks for downstream consumers.

I would expect this operation on the catalog itself:

catalog=load_catalog('default')
catalog.create_table('schema.table', schema=...)
catalog.create_or_replace_table('schema.table', schema=...)

We want to generalize this operation, so we don't have to implement this for each of the catalogs. Therefore I would expect this on the Catalog(ABC) itself.

Just a heads up, for the replace table it keeps the history in Spark:

image

And when we look at the metadata, we can see the previous schema/snapshot as well:

{
"format-version" : 2,
"table-uuid" : "9b8b02af-2097-453f-86e2-5b2715e9d37a",
"location" : "s3://warehouse/default/fokko",
"last-sequence-number" : 2,
"last-updated-ms" : 1708081058809,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age""required" : false,
"type" : "int"
} ],
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-02-16T10:57:38.541088095Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : 398515508184271470,
"refs" : {
"main" : {
"snapshot-id" : 398515508184271470,
"type" : "branch"
}
},
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 4615041670163082108,
"timestamp-ms" : 1708081058629,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-4615041670163082108-1-d3852ba7-ff54-4abd-99a2-0265206cfbfa.avro",
"schema-id" : 0
}, {
"sequence-number" : 2,
"snapshot-id" : 398515508184271470,
"timestamp-ms" : 1708081058809,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1708080918556",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "628",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "628",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/fokko/metadata/snap-398515508184271470-1-4d03a8b5-8912-4235-8c18-75400fef9874.avro",
"schema-id" : 1
} ],
"statistics" : [ ],
"snapshot-log" : [ {
"timestamp-ms" : 1708081058809,
"snapshot-id" : 398515508184271470
} ],
"metadata-log" : [ {
"timestamp-ms" : 1708081058629,
"metadata-file" : "s3://warehouse/default/fokko/metadata/00000-10d2c8d5-f6a2-4dc4-90cb-c545d8ffd497.metadata.json"
} ]
}

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Thank you @Fokko for taking time to explain in such great detail. Now it makes much more sense to have this part of the Catalog API. Made changes as suggested.

@FokkoFokko added this to the PyIceberg 0.7.0 release milestone Feb 19, 2024
@anupam-saini
anupam-saini marked this pull request as ready for review February 21, 2024 03:54
@sungwysungwy mentioned this pull request Feb 22, 2024
@anupam-saini
anupam-saini requested a review from sungwyMarch 1, 2024 03:12
@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

Now with Sort Order and Partition Spec updates, this PR has all the necessary pieces for create-replace table operation and is ready for review.

@Fokko @syun64

@FokkoFokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the late reply @anupam-saini This is looking great, but looks like there are still some discrepancies with the reference implementation in Java.

AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),
)
tx._apply(updates, requirements)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With #471 in this is not necessary anymore! 🥳

Suggested change
tx._apply(updates, requirements)
tx._apply(updates, requirements)

table = self.load_table(identifier)
with table.transaction() as tx:
base_schema = table.schema()
new_schema = assign_fresh_schema_ids(schema_or_type=new_schema, base_schema=base_schema)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's consider the following:

CREATETABLEdefault.t1 (name string);

Results in:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407936565,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t1 (name string, age int);

The second schema is added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710407992389,
"last-column-id" : 2,
"current-schema-id" : 1,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
} ]
}

And then go back to the original schema:

CREATE OR REPLACETABLEdefault.t1 (name string);

You'll see that no new schema is being added:

{
"format-version" : 2,
"table-uuid" : "c10da17c-c40c-4e02-9e81-15b6f0f35cb9",
"location" : "s3://warehouse/default/t1",
"last-sequence-number" : 0,
"last-updated-ms" : 1710408026710,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
}, {
"type" : "struct",
"schema-id" : 1,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710407936565,
"metadata-file" : "s3://warehouse/default/t1/metadata/00000-83e6fd53-d3c9-41f9-aeea-b978fb882c7f.metadata.json"
}, {
"timestamp-ms" : 1710407992389,
"metadata-file" : "s3://warehouse/default/t1/metadata/00001-00192fa8-9019-481a-b4da-ebd99d11eeb9.metadata.json"
} ]
}

What do you think of re-using the update_schema() class:

withtable.transaction() astransaction:
withtransaction.update_schema(allow_incompatible_changes=True) asupdate_schema:
# Remove old fieldsremoved_column_names=base_schema._name_to_id().keys() -schema._name_to_id().keys()
forremoved_column_nameinremoved_column_names:
update_schema.delete_column(removed_column_name)
# Add new and evolve existing fieldsupdate_schema.union_by_name(schema)

^ Pseudocode, could be cleaner. Ideally, the removal should be done with a visit_with_partner (that's the opposite of the union_by_name.

@anupam-sainianupam-sainiMar 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Fokko for this detailed explanation.
But we also need to cover the step 2 of this example where we add a new schema, right?

So from my understanding, if the schema fields match with an old schema in the metadata, we do union_by_name with the old schema and set it as the current one
Else, we add the new schema.
Is this correct assessment?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct!

Comment on lines +792 to +793
AddPartitionSpecUpdate(spec=new_partition_spec),
SetDefaultSpecUpdate(spec_id=-1),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same goes here, the spec is being re-used:

CREATETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409060360,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
} ],
"last-partition-id" : 1000,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name, age);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409079414,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 1,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t2 (name string, age int) PARTITIONED BY (name);
{
"format-version" : 2,
"table-uuid" : "27f1ab29-7a9f-4324-bb64-c10c5bd2be53",
"location" : "s3://warehouse/default/t2",
"last-sequence-number" : 0,
"last-updated-ms" : 1710409086268,
"last-column-id" : 2,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
}, {
"id" : 2,
"name" : "age",
"required" : false,
"type" : "int"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
} ]
}, {
"spec-id" : 1,
"fields" : [ {
"name" : "name",
"transform" : "identity",
"source-id" : 1,
"field-id" : 1000
}, {
"name" : "age",
"transform" : "identity",
"source-id" : 2,
"field-id" : 1001
} ]
} ],
"last-partition-id" : 1001,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409060360,
"metadata-file" : "s3://warehouse/default/t2/metadata/00000-bc21fe27-d037-4b98-a8ca-cc1502eb19ec.metadata.json"
}, {
"timestamp-ms" : 1710409079414,
"metadata-file" : "s3://warehouse/default/t2/metadata/00001-8062294f-a8d6-493d-905f-b82dfe01cb29.metadata.json"
} ]
}

Ideally, we also want to-reuse the update_spec class here.

)

requirements: Tuple[TableRequirement, ...] = (AssertTableUUID(uuid=table.metadata.table_uuid),)
updates: Tuple[TableUpdate, ...] = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need to clear the snapshots here as well:

CREATETABLEdefault.t3 ASSELECT'Fokko'as name
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710409653861,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
} ]
}
CREATE OR REPLACETABLEdefault.t3 (name string);
{
"format-version" : 2,
"table-uuid" : "c61404c8-2211-46c7-866f-2eb87022b728",
"location" : "s3://warehouse/default/t3",
"last-sequence-number" : 1,
"last-updated-ms" : 1710411760623,
"last-column-id" : 1,
"current-schema-id" : 0,
"schemas" : [ {
"type" : "struct",
"schema-id" : 0,
"fields" : [ {
"id" : 1,
"name" : "name",
"required" : false,
"type" : "string"
} ]
} ],
"default-spec-id" : 0,
"partition-specs" : [ {
"spec-id" : 0,
"fields" : [ ]
} ],
"last-partition-id" : 999,
"default-sort-order-id" : 0,
"sort-orders" : [ {
"order-id" : 0,
"fields" : [ ]
} ],
"properties" : {
"owner" : "root",
"created-at" : "2024-03-14T09:47:10.455199504Z",
"write.parquet.compression-codec" : "zstd"
},
"current-snapshot-id" : -1,
"refs" : { },
"snapshots" : [ {
"sequence-number" : 1,
"snapshot-id" : 3622792816294171432,
"timestamp-ms" : 1710409631964,
"summary" : {
"operation" : "append",
"spark.app.id" : "local-1710405058122",
"added-data-files" : "1",
"added-records" : "1",
"added-files-size" : "416",
"changed-partition-count" : "1",
"total-records" : "1",
"total-files-size" : "416",
"total-data-files" : "1",
"total-delete-files" : "0",
"total-position-deletes" : "0",
"total-equality-deletes" : "0"
},
"manifest-list" : "s3://warehouse/default/t3/metadata/snap-3622792816294171432-1-e457c732-62e5-41eb-998b-abbb8f021ed5.avro",
"schema-id" : 0
} ],
"statistics" : [ ],
"snapshot-log" : [ ],
"metadata-log" : [ {
"timestamp-ms" : 1710409631964,
"metadata-file" : "s3://warehouse/default/t3/metadata/00000-c82191f3-e6e2-4001-8e85-8623e3915ff7.metadata.json"
}, {
"timestamp-ms" : 1710409653861,
"metadata-file" : "s3://warehouse/default/t3/metadata/00001-0297d4e7-2468-4c0d-b4ed-ea717df8c3e6.metadata.json"
} ]
}

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@anupam-saini Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

Feel free to take over. The initial approach we had in mind had diverged from the overall direction.

@anupam-saini

Copy link
Copy Markdown
ContributorAuthor

@srilman let me know if you would like to open a new PR. I can close this one

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor

Are you still working on this PR? For context, we are doing some interesting (and probably wrong) stuff to mimic Iceberg Java's REPLACE TABLE in https://github.com/bodo-ai/Bodo/blob/main/bodo/io/iceberg/write.py#L610, so having this built-in would be useful. If you're not, happy to take over the PR and address the comments

+1 from me - thank you!

@srilman

Copy link
Copy Markdown
Contributor

@anupam-saini Yep was planning to. Feel free to close this one

@Fokko

Fokko commented Jun 1, 2025

Copy link
Copy Markdown
Contributor

Closing this as suggested by @anupam-saini. Looking forward seeing this being added to PyIceberg 🚀

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

REPLACE TABLE Support

7 participants

@anupam-saini@Fokko@srilman@smaheshwar-pltr@sungwy@kevinjqliu@HonahX