Skip to content

Spark 4.1: Implement RepairTable action - #17622

Open
rahulsmahadev wants to merge 6 commits into
apache:mainfrom
rahulsmahadev:repair-table-spark-impl
Open

Spark 4.1: Implement RepairTable action#17622
rahulsmahadev wants to merge 6 commits into
apache:mainfrom
rahulsmahadev:repair-table-spark-impl

Conversation

@rahulsmahadev

Copy link
Copy Markdown
Contributor

Spark implementation of the RepairTable action whose API was added in #17399.

repairFileMetrics() compares the metrics recorded in each live manifest entry (record count, file size, column bounds/null/nan/value counts) against the underlying data and delete files, and rewrites only the manifests that contain an incorrect entry. Rewritten entries carry through their original snapshot id and sequence numbers, so the repair does not change which delete files apply. dryRun() reports what would be repaired without committing.

Scoped to Spark 4.1 for now; happy to backport once this lands.

Tests in TestRepairTableAction cover: no-op on correct stats, record-count/file-size/column-metric repair, dry run, partitioned tables, and concurrency (concurrent append, a conflicting manifest rewrite, commit-state-unknown, and cleanup of manifests written by a failed or dry-run repair).

@rahulsmahadev
rahulsmahadev marked this pull request as ready for review August 17, 2026 18:35
@rahulsmahadev

Copy link
Copy Markdown
ContributorAuthor

return builder
.withFileSizeInBytes(corruptCounts ? file.fileSizeInBytes() + 4096 : file.fileSizeInBytes())
.build();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TODO: add test with custom stats column (not 100) and contains the iceberg property for stats column

@anuragmantrianuragmantri left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @rahulsmahadev. This is great and solves a repeated pain point for our customers with corrupted data.

I did an initial review and found that partition evolution causes the logic to write incorrect spec ids. Details in the comments below.

return new ManifestWriterFactory(
sparkContext().broadcast(SerializableTableWithSize.copyOf(table)),
formatVersion,
table.spec().specId(),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Load manifests in L301 reads every data/delete manifest of the current snapshot regardless of partitionSpecId() and here in the writer they are bound to table.spec().specId(), which is the table's current default spec.

If there was partition evolution and we are reading older manifests, this will write incorrect partition ids to older schemas. This is the same class of bug as #666

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I ran this test and verified it is indeed an issue

@TestTemplatepublicvoidtestRepairAfterPartitionSpecEvolution() throwsIOException {
Tabletable = createTable(PartitionSpec.unpartitioned());
appendRecords(table, records(4));
DataFileoriginal = onlyDataFile(table);
assertThat(original.specId()).isEqualTo(0);
assertThat(original.partition().size()).isEqualTo(0);
// evolve the table to a partitioned spec; the existing manifest keeps referring to spec 0table.updateSpec().addField("c1").commit();
table.refresh();
assertThat(table.spec().specId()).isEqualTo(1);
ManifestFileoldManifest = table.currentSnapshot().dataManifests(table.io()).get(0);
assertThat(oldManifest.partitionSpecId())
.as("the manifest written before the evolution must still be tagged with the old spec")
.isEqualTo(0);
// corrupt the stats of the entry that still belongs to the original, unpartitioned speccorruptStats(table, oldManifest, original.location());
SparkActions.get().repairTable(table).execute();
table.refresh();
DataFilerepaired = onlyDataFile(table);
assertThat(repaired.recordCount())
.as("the repair must still correct the stats")
.isEqualTo(original.recordCount());
assertThat(repaired.specId())
.as("the repaired entry must keep the spec it was originally written under")
.isEqualTo(0);
assertThat(repaired.partition().size())
.as("an unpartitioned file's partition data must still have zero fields after repair")
.isEqualTo(0);
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very good catch and thanks for the repro, handled it in 063328f

? new WriteDataManifests(writers, combinedFileType, sparkType, repaired, context)
: new WriteDeleteManifests(writers, combinedFileType, sparkType, repaired, context);

// preserve the entry order of the manifests being rewritten

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did not understand this comment. I don't think repartition(n) preserves order. Can you clarify?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated the comment, you are right

@szehon-hoszehon-ho left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The data file path mirrors RewriteManifestsSparkAction closely and the concurrency tests are thorough. My main concern is the delete file path: two issues there look like real correctness bugs, and that path currently has no test coverage. Details inline.


Reviewed with AI assistance (Claude Opus 5 via Cursor). Findings were checked against the source by hand. The two correctness findings were derived from reading the builders and DeleteFilter, not from a reproduction.

/** Returns the name mapping of the table, or null if the table does not define one. */
static NameMapping nameMapping(Table table) {
String mapping =
table.properties().get(org.apache.iceberg.TableProperties.DEFAULT_NAME_MAPPING);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Import DEFAULT_NAME_MAPPING rather than qualifying it inline. Same for org.apache.iceberg.io.InputFile at RepairTableSparkAction L562 and L704, and java.util.function.Function at L388.

.withFileSizeInBytes(fileSizeInBytes)
.build();
} else {
return FileMetadata.deleteFileBuilder(spec)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add .ofEqualityDeletes(...) when the file is an equality delete. FileMetadata.Builder.copy(DeleteFile) doesn't carry equalityFieldIds and build() doesn't validate them, so the rebuilt entry keeps content = EQUALITY_DELETES but writes equality_ids = null. DeleteFilter then does Sets.newHashSet(delete.equalityFieldIds()) and NPEs on the next read, with the bad entry already committed.

RewriteTablePathUtil.newEqualityDeleteEntry restores the ids explicitly for this reason.

*/
public static final String REPAIR_COLUMN_METRICS = "repair-column-metrics";

public static final boolean REPAIR_COLUMN_METRICS_DEFAULT = true;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reword the javadoc above: disabling this doesn't avoid reading footers, since CheckStats calls readMetrics unconditionally and the option only skips the comparison.

Separately, consider defaulting to false, or restricting the comparison to columns present in both maps. MetricsConfig.forTable reflects the table's current schema and sort order and nothing records the config a file was written under, so true will rewrite manifests over stats that were correct when written. The clearest case is column count: once a table grows past write.metadata.metrics.max-inferred-column-defaults (100), MetricsConfig.from assigns the default mode to only the first 100 field ids and None to the rest, so recomputed metrics carry fewer bounds than entries written while the table was narrower. boundsMatch fails, the entry is flagged, and withStats writes the reduced set — the repair drops valid bounds. Setting write.metadata.metrics.default to none or counts gets there too, via the sorted-column promotion.

}

@Override
public RepairTableSparkAction repairFileMetrics() {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Set a field here and gate the repair on it, so execute() is a no-op when no repair has been selected. RepairTable's javadoc says the repairs are selected through the configuration methods, but today the flag is unobservable and execute() repairs metrics regardless.

return EMPTY_RESULT;
}

if (dryRun) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Return before writeManifests on a dry run rather than writing manifests and deleting them here. Both result fields come from the verdicts, so the write isn't needed, and a failure between writing and this cleanup leaves orphan repaired-m-*.avro files in the metadata directory. If the write is deliberate as a rehearsal of the real path, a note on dryRun would help, since creating files is surprising for a dry run.

}

SparkContentFile<?> newFileWrapper(Types.StructType combinedFileType, StructType sparkType) {
Types.StructType fileType = DataFile.getType(table().spec().partitionType());

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pass the group's spec here, or null to skip the projection since it's unused. This projects to the current default spec, the same thing behind the bug @anuragmantri caught on the write side; it's harmless today only because CheckStats never reads partition().

}

/** A manifest entry whose statistics disagree with the file it refers to. */
public static class EntryVerdict implements Serializable {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use Encoders.tuple(Encoders.STRING(), Encoders.STRING()), or select the two columns and collect Rows. Either avoids adding a public bean with setters to org.apache.iceberg.spark.actions just to satisfy Encoders.bean.

Row fileRow = row.getStruct(4);
StructType sparkType = (StructType) fileRow.schema();
Types.StructType combinedFileType =
DataFile.getType(Partitioning.partitionType(context.table()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hoist combinedFileType and the wrapper out of the loop, since they're identical for every row. As written each entry recomputes Partitioning.partitionType(table), rebuilds the DataFile struct type, and constructs a SparkContentFile, which runs StructProjection.create and builds a field position map. WriteManifests.call creates its wrapper once per partition.

Metrics metrics =
RepairMetrics.readMetrics(
input, file, context.metricsConfig(file), context.nameMapping());
return RepairMetrics.withStats(file, context.spec(file.specId()), metrics, fileSizeInBytes);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Preserve the existing column stats when repair-column-metrics is false. repairStats applies the full recomputed Metrics regardless, so an entry flagged only for a wrong record count or file size still gets its column stats replaced — exactly the case the option exists to protect against, since its rationale is that a table whose metrics config changed reports legitimately different column stats.


@Parameters(name = "formatVersion = {0}")
public static Object[] parameters() {
return new Object[][] {new Object[] {1}, new Object[] {2}, new Object[] {3}};

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add coverage for delete manifests: a position delete with wrong stats, an equality delete with wrong stats, a manifest holding both, and a v3 manifest with a DV next to a repairable entry asserting referencedDataFile, contentOffset and contentSizeInBytes survive. These cases are all data files, so WriteDeleteManifests and SparkDeleteFile ship unexercised even though doExecute iterates ManifestContent.values(). The first two would catch the issues I flagged in RepairMetrics.withStats and RepairContext.metricsConfig.

Adds a Spark implementation of the RepairTable action, which repairs
manifest entries whose statistics disagree with the files they refer to.
The statistics of every live entry are compared against the file by reading
its footer, and only the manifests that contain at least one incorrect entry
are rewritten, so the cost of the commit is proportional to the number of
incorrect entries rather than to the size of the table.
Entries are rewritten with ManifestWriter#existing, carrying through the
original snapshot id and data and file sequence numbers. This preserves the
lineage of the files, and therefore which delete files apply to them, so the
repair leaves the contents of the table unchanged.
Repairing statistics does not change the number of live files, so the commit
goes through the existing rewrite manifests validation unchanged.
- RepairMetrics reads and compares statistics per format (Parquet, ORC and
Avro) and rebuilds a file with corrected statistics.
- Files whose statistics cannot be read are carried through unchanged and
counted as incorrect but not repaired, which is what distinguishes
entryStatsIncorrectCount from entryStatsRepairedCount.
- The repair-column-metrics option skips the footer reads and repairs only
record counts and file sizes, for tables where only those are suspect.
- dryRun reports what would be repaired without committing.
Covers what happens when the table changes underneath a repair, and what is
left behind when the commit does not succeed:
- a concurrent append commits between planning and the repair commit: the
repair succeeds and the appended records survive, since the appended data
lands in a new manifest and the manifests being repaired are still present
- a concurrent operation replaces the very manifest being repaired: the
commit fails validation in BaseRewriteManifests#validateDeletedManifests
and the table is left untouched, rather than dropping the concurrent change
- a failed commit deletes the manifests the repair wrote
- a commit reported as CommitStateUnknownException keeps them, as the commit
may have succeeded
- a dry run leaves none of them behind
Note the cleanup tests assert on the manifests written by the action itself.
A failed commit on a format version 1 table can also leave behind a copy of
a manifest made by the core staging path in BaseRewriteManifests, which is
only cleaned up after a successful commit and is not owned by this action.
Manifests of the current snapshot were loaded regardless of their partition
spec, but the writer was bound to the current default spec of the table. After
partition evolution, rewriting a manifest of an older spec therefore wrote the
current spec id and partition type to its entries.
Group the manifests by partition spec id and repair each group with a writer
bound to that spec, so an entry keeps the spec it was written under. Also pass
the spec specific file type to the writer instead of deriving it from the
current spec of the table.
Correctness:
- Equality deletes now keep their equality field ids when rebuilt.
FileMetadata.Builder.copy(DeleteFile) drops them, so a repaired entry
otherwise kept content EQUALITY_DELETES with null equality ids, and
reading the table failed once the delete was applied.
- The metrics config is resolved per file content, so a delete manifest
holding both position and equality deletes no longer compares equality
deletes against the position-delete config and rewrites them needlessly.
- The check-side file wrapper binds to the manifest's own spec rather than
the table's current spec.
- When column-metrics repair is disabled, only the record count and file
size of a flagged entry are corrected; the stored column stats are kept
instead of being replaced with the recomputed ones.
Behavior:
- repairFileMetrics() now selects the repair; execute() is a no-op when no
repair has been selected, as the interface describes.
- repair-column-metrics defaults to false. Recomputed column stats reflect
the current metrics config, which the file may not have been written
under, so repairing them can overwrite correct statistics.
- A dry run no longer writes manifests only to delete them; it reports what
would be repaired without writing anything.
Also hoist the combined file type and file wrapper out of the per-row loop
in CheckStats, and replace inline-qualified references with imports.
… delete manifests
The set of incorrect file paths was collected to the driver and broadcast back
out. That set is unbounded, since a writer that recorded stats incorrectly
usually did so for every file it wrote. Mark the entries to repair by joining
the entries against the verdicts on the file path instead, so the set stays
distributed. Only the set of manifests to rewrite, which is naturally small, is
still collected. Verdicts are emitted as tuples rather than a bean and cached,
as they are read more than once.
Repartition the entries by manifest when writing so the manifest layout of the
table is preserved, rather than scattered round robin by a plain repartition(n).
Guard the equality field ids in RepairMetrics.withStats: an entry that is an
equality delete but records no equality ids is carried through as is rather than
throwing, so a malformed entry does not abort the whole repair.
Add end-to-end coverage of the delete manifest path, which was previously
unexercised: repair a position delete, an equality delete, and a manifest
holding both, asserting the statistics are corrected and the equality field ids
survive.
The two tests asserting that a failed repair leaves the table unchanged scan the table while the corruption is still present. Parquet now fetches small files eagerly using the recorded file size, so an inflated file_size_in_bytes makes the scan read past the end of the file. Corrupt only the record count in these tests, keeping the file size accurate, so the table stays readable.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@rahulsmahadev@anuragmantri@szehon-ho