Skip to content

[Python][Parquet] read_schema drops extension types (UUID returned as fixed_size_binary[16]) #48254

Description

@Kuinox

Summary

UUID extension types are preserved in tables but dropped by pyarrow.parquet.read_schema, creating an asymmetry between the table’s schema and the schema read from Parquet metadata.

Steps to Reproduce

importpyarrowaspaimportpyarrow.parquetaspqfrompathlibimportPathimporttempfiledata= [
b'\xe4`\xf9p\x83QGN\xac\x7f\xa4g>\x4b\xa8\xcb',
b'\x1et\x14\x95\xee\xd5C\xea\x9b\xd7s\xdc\x91BK\xaf',
None,
]
table=pa.table([pa.array(data, type=pa.uuid())], names=["ext"])
print("table schema type:", table.schema.field("ext").type) # extension<arrow.uuid>path=Path(tempfile.gettempdir()) /"uuid_ext_test.parquet"pq.write_table(table, path, store_schema=False)
print("read_schema type:", pq.read_schema(path).field("ext").type)
print("read_table schema type:", pq.read_table(path).schema.field("ext").type)

Expected Behavior

read_schema(path) should yield the same type as the table schema (and read_table), i.e., extension<arrow.uuid>.

Actual Behavior

read_schema(path) returns fixed_size_binary[16], while the original table.schema and read_table(path).schema both report extension<arrow.uuid>, so metadata-based schema inspection drops the extension type.

Notes

  • Observed with the current pyarrow wheel (22.0.0) and current main sources.
  • ParquetFile(...).schema_arrow preserves the extension type; read_schema does not

Component(s)

Python, Parquet

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions