Null list<struct<...>> is written and read as an empty listΒ #3833

Description

@slachiewicz

Apache Iceberg version

main (development)

Please describe the bug 🐞

A null list<struct<...>> is silently rebuilt as an empty list. This is known on
the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
asserts the corrupted value, with the correct assertion commented out pending
apache/arrow#38809:

# This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

Two things seem worth reporting on top of that.

It also affects the write path, where the consequence is worse. The Parquet file
pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
this particular null loss can be fixed independently of apache/arrow#38809, which is
still open.

This is the array<struct<>> case from #251. That issue was closed in March 2025 on
the strength of this test existing, but the assertion it makes is the corrupted one;
the array<int> case in the issue body was genuinely fixed by #252, while the
array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
remaining broken in
#252 (comment) β€” was not.

Reproduction (write path)

pyiceberg 0.11.1, pyarrow 25.0.1:

importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
sch=pa.schema([
pa.field("id", pa.int32(), nullable=False),
pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
])
tbl=pa.table({"id": [1, 2, 3, 4],
"l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
"l_int": [[1], [], None, [3]]}, schema=sch)
cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
cat.create_namespace("ns")
it=cat.create_table("ns.t", schema=tbl.schema)
it.append(tbl)
out=it.scan().to_arrow()
forcin ("l_struct", "l_int"):
print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
# the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

Output:

l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]

l_int round-trips correctly, and writing the same pa.Table with pq.write_table
preserves the null, so the loss is not pyarrow's.

Cause

ArrowProjectionVisitor.list rebuilds the array when the element is a struct
(pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

ifisinstance(value_array, pa.StructArray):
# This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

from_arrays receives the offsets buffer alone, which cannot express a null list, and
no mask, so the validity bitmap is dropped. That is also why only this one shape is
affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
does not rebuild at all, and a list whose element is a primitive never enters this
branch. The visitor runs on both paths, which is why the same root cause shows up as
the read-side assertion above and as the write-side corruption here.

Fix

Carrying the mask over is enough:

list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

With that change the reproduction above returns None for both columns, and
test_null_list_and_map passes with its commented-out assertion restored. I have not
looked at int32-offset or sliced-array handling of this call, which the existing line
already relies on; that appears independent of the mask.

I have this on a branch with a unit test covering the write path and the integration
assertion un-commented, and can open a PR.

Found while testing a third-party Iceberg writer against pyiceberg as a reader.

Willingness to contribute

  • I can contribute a fix for this bug independently

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      Null list<struct<...>> is written and read as an empty listΒ #3833

      Description

      @slachiewicz

      Apache Iceberg version

      main (development)

      Please describe the bug 🐞

      A null list<struct<...>> is silently rebuilt as an empty list. This is known on
      the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
      asserts the corrupted value, with the correct assertion commented out pending
      apache/arrow#38809:

      # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

      Two things seem worth reporting on top of that.

      It also affects the write path, where the consequence is worse. The Parquet file
      pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
      pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

      It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
      takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
      this particular null loss can be fixed independently of apache/arrow#38809, which is
      still open.

      This is the array<struct<>> case from #251. That issue was closed in March 2025 on
      the strength of this test existing, but the assertion it makes is the corrupted one;
      the array<int> case in the issue body was genuinely fixed by #252, while the
      array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
      remaining broken in
      #252 (comment) β€” was not.

      Reproduction (write path)

      pyiceberg 0.11.1, pyarrow 25.0.1:

      importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
      sch=pa.schema([
      pa.field("id", pa.int32(), nullable=False),
      pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
      pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
      ])
      tbl=pa.table({"id": [1, 2, 3, 4],
      "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
      "l_int": [[1], [], None, [3]]}, schema=sch)
      cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
      cat.create_namespace("ns")
      it=cat.create_table("ns.t", schema=tbl.schema)
      it.append(tbl)
      out=it.scan().to_arrow()
      forcin ("l_struct", "l_int"):
      print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
      # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
      print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

      Output:

      l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
      l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
      raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
      

      l_int round-trips correctly, and writing the same pa.Table with pq.write_table
      preserves the null, so the loss is not pyarrow's.

      Cause

      ArrowProjectionVisitor.list rebuilds the array when the element is a struct
      (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

      ifisinstance(value_array, pa.StructArray):
      # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

      from_arrays receives the offsets buffer alone, which cannot express a null list, and
      no mask, so the validity bitmap is dropped. That is also why only this one shape is
      affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
      does not rebuild at all, and a list whose element is a primitive never enters this
      branch. The visitor runs on both paths, which is why the same root cause shows up as
      the read-side assertion above and as the write-side corruption here.

      Fix

      Carrying the mask over is enough:

      list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

      With that change the reproduction above returns None for both columns, and
      test_null_list_and_map passes with its commented-out assertion restored. I have not
      looked at int32-offset or sliced-array handling of this call, which the existing line
      already relies on; that appears independent of the mask.

      I have this on a branch with a unit test covering the write path and the integration
      assertion un-commented, and can open a PR.

      Found while testing a third-party Iceberg writer against pyiceberg as a reader.

      Willingness to contribute

      • I can contribute a fix for this bug independently

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          Null list<struct<...>> is written and read as an empty listΒ #3833

          Description

          @slachiewicz

          Apache Iceberg version

          main (development)

          Please describe the bug 🐞

          A null list<struct<...>> is silently rebuilt as an empty list. This is known on
          the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
          asserts the corrupted value, with the correct assertion commented out pending
          apache/arrow#38809:

          # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

          Two things seem worth reporting on top of that.

          It also affects the write path, where the consequence is worse. The Parquet file
          pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
          pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

          It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
          takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
          this particular null loss can be fixed independently of apache/arrow#38809, which is
          still open.

          This is the array<struct<>> case from #251. That issue was closed in March 2025 on
          the strength of this test existing, but the assertion it makes is the corrupted one;
          the array<int> case in the issue body was genuinely fixed by #252, while the
          array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
          remaining broken in
          #252 (comment) β€” was not.

          Reproduction (write path)

          pyiceberg 0.11.1, pyarrow 25.0.1:

          importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
          sch=pa.schema([
          pa.field("id", pa.int32(), nullable=False),
          pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
          pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
          ])
          tbl=pa.table({"id": [1, 2, 3, 4],
          "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
          "l_int": [[1], [], None, [3]]}, schema=sch)
          cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
          cat.create_namespace("ns")
          it=cat.create_table("ns.t", schema=tbl.schema)
          it.append(tbl)
          out=it.scan().to_arrow()
          forcin ("l_struct", "l_int"):
          print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
          # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
          print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

          Output:

          l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
          l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
          raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
          

          l_int round-trips correctly, and writing the same pa.Table with pq.write_table
          preserves the null, so the loss is not pyarrow's.

          Cause

          ArrowProjectionVisitor.list rebuilds the array when the element is a struct
          (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

          ifisinstance(value_array, pa.StructArray):
          # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

          from_arrays receives the offsets buffer alone, which cannot express a null list, and
          no mask, so the validity bitmap is dropped. That is also why only this one shape is
          affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
          does not rebuild at all, and a list whose element is a primitive never enters this
          branch. The visitor runs on both paths, which is why the same root cause shows up as
          the read-side assertion above and as the write-side corruption here.

          Fix

          Carrying the mask over is enough:

          list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

          With that change the reproduction above returns None for both columns, and
          test_null_list_and_map passes with its commented-out assertion restored. I have not
          looked at int32-offset or sliced-array handling of this call, which the existing line
          already relies on; that appears independent of the mask.

          I have this on a branch with a unit test covering the write path and the integration
          assertion un-commented, and can open a PR.

          Found while testing a third-party Iceberg writer against pyiceberg as a reader.

          Willingness to contribute

          • I can contribute a fix for this bug independently

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Null list<struct<...>> is written and read as an empty listΒ #3833

              Description

              @slachiewicz

              Apache Iceberg version

              main (development)

              Please describe the bug 🐞

              A null list<struct<...>> is silently rebuilt as an empty list. This is known on
              the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
              asserts the corrupted value, with the correct assertion commented out pending
              apache/arrow#38809:

              # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

              Two things seem worth reporting on top of that.

              It also affects the write path, where the consequence is worse. The Parquet file
              pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
              pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

              It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
              takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
              this particular null loss can be fixed independently of apache/arrow#38809, which is
              still open.

              This is the array<struct<>> case from #251. That issue was closed in March 2025 on
              the strength of this test existing, but the assertion it makes is the corrupted one;
              the array<int> case in the issue body was genuinely fixed by #252, while the
              array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
              remaining broken in
              #252 (comment) β€” was not.

              Reproduction (write path)

              pyiceberg 0.11.1, pyarrow 25.0.1:

              importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
              sch=pa.schema([
              pa.field("id", pa.int32(), nullable=False),
              pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
              pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
              ])
              tbl=pa.table({"id": [1, 2, 3, 4],
              "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
              "l_int": [[1], [], None, [3]]}, schema=sch)
              cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
              cat.create_namespace("ns")
              it=cat.create_table("ns.t", schema=tbl.schema)
              it.append(tbl)
              out=it.scan().to_arrow()
              forcin ("l_struct", "l_int"):
              print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
              # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
              print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

              Output:

              l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
              l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
              raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
              

              l_int round-trips correctly, and writing the same pa.Table with pq.write_table
              preserves the null, so the loss is not pyarrow's.

              Cause

              ArrowProjectionVisitor.list rebuilds the array when the element is a struct
              (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

              ifisinstance(value_array, pa.StructArray):
              # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

              from_arrays receives the offsets buffer alone, which cannot express a null list, and
              no mask, so the validity bitmap is dropped. That is also why only this one shape is
              affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
              does not rebuild at all, and a list whose element is a primitive never enters this
              branch. The visitor runs on both paths, which is why the same root cause shows up as
              the read-side assertion above and as the write-side corruption here.

              Fix

              Carrying the mask over is enough:

              list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

              With that change the reproduction above returns None for both columns, and
              test_null_list_and_map passes with its commented-out assertion restored. I have not
              looked at int32-offset or sliced-array handling of this call, which the existing line
              already relies on; that appears independent of the mask.

              I have this on a branch with a unit test covering the write path and the integration
              assertion un-commented, and can open a PR.

              Found while testing a third-party Iceberg writer against pyiceberg as a reader.

              Willingness to contribute

              • I can contribute a fix for this bug independently

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  Null list<struct<...>> is written and read as an empty listΒ #3833

                  Description

                  @slachiewicz

                  Apache Iceberg version

                  main (development)

                  Please describe the bug 🐞

                  A null list<struct<...>> is silently rebuilt as an empty list. This is known on
                  the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
                  asserts the corrupted value, with the correct assertion commented out pending
                  apache/arrow#38809:

                  # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

                  Two things seem worth reporting on top of that.

                  It also affects the write path, where the consequence is worse. The Parquet file
                  pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
                  pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

                  It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
                  takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
                  this particular null loss can be fixed independently of apache/arrow#38809, which is
                  still open.

                  This is the array<struct<>> case from #251. That issue was closed in March 2025 on
                  the strength of this test existing, but the assertion it makes is the corrupted one;
                  the array<int> case in the issue body was genuinely fixed by #252, while the
                  array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
                  remaining broken in
                  #252 (comment) β€” was not.

                  Reproduction (write path)

                  pyiceberg 0.11.1, pyarrow 25.0.1:

                  importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
                  sch=pa.schema([
                  pa.field("id", pa.int32(), nullable=False),
                  pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
                  pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
                  ])
                  tbl=pa.table({"id": [1, 2, 3, 4],
                  "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
                  "l_int": [[1], [], None, [3]]}, schema=sch)
                  cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
                  cat.create_namespace("ns")
                  it=cat.create_table("ns.t", schema=tbl.schema)
                  it.append(tbl)
                  out=it.scan().to_arrow()
                  forcin ("l_struct", "l_int"):
                  print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
                  # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
                  print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

                  Output:

                  l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
                  l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
                  raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
                  

                  l_int round-trips correctly, and writing the same pa.Table with pq.write_table
                  preserves the null, so the loss is not pyarrow's.

                  Cause

                  ArrowProjectionVisitor.list rebuilds the array when the element is a struct
                  (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

                  ifisinstance(value_array, pa.StructArray):
                  # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

                  from_arrays receives the offsets buffer alone, which cannot express a null list, and
                  no mask, so the validity bitmap is dropped. That is also why only this one shape is
                  affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
                  does not rebuild at all, and a list whose element is a primitive never enters this
                  branch. The visitor runs on both paths, which is why the same root cause shows up as
                  the read-side assertion above and as the write-side corruption here.

                  Fix

                  Carrying the mask over is enough:

                  list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

                  With that change the reproduction above returns None for both columns, and
                  test_null_list_and_map passes with its commented-out assertion restored. I have not
                  looked at int32-offset or sliced-array handling of this call, which the existing line
                  already relies on; that appears independent of the mask.

                  I have this on a branch with a unit test covering the write path and the integration
                  assertion un-commented, and can open a PR.

                  Found while testing a third-party Iceberg writer against pyiceberg as a reader.

                  Willingness to contribute

                  • I can contribute a fix for this bug independently

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      Null list<struct<...>> is written and read as an empty listΒ #3833

                      Description

                      @slachiewicz

                      Apache Iceberg version

                      main (development)

                      Please describe the bug 🐞

                      A null list<struct<...>> is silently rebuilt as an empty list. This is known on
                      the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
                      asserts the corrupted value, with the correct assertion commented out pending
                      apache/arrow#38809:

                      # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

                      Two things seem worth reporting on top of that.

                      It also affects the write path, where the consequence is worse. The Parquet file
                      pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
                      pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

                      It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
                      takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
                      this particular null loss can be fixed independently of apache/arrow#38809, which is
                      still open.

                      This is the array<struct<>> case from #251. That issue was closed in March 2025 on
                      the strength of this test existing, but the assertion it makes is the corrupted one;
                      the array<int> case in the issue body was genuinely fixed by #252, while the
                      array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
                      remaining broken in
                      #252 (comment) β€” was not.

                      Reproduction (write path)

                      pyiceberg 0.11.1, pyarrow 25.0.1:

                      importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
                      sch=pa.schema([
                      pa.field("id", pa.int32(), nullable=False),
                      pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
                      pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
                      ])
                      tbl=pa.table({"id": [1, 2, 3, 4],
                      "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
                      "l_int": [[1], [], None, [3]]}, schema=sch)
                      cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
                      cat.create_namespace("ns")
                      it=cat.create_table("ns.t", schema=tbl.schema)
                      it.append(tbl)
                      out=it.scan().to_arrow()
                      forcin ("l_struct", "l_int"):
                      print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
                      # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
                      print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

                      Output:

                      l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
                      l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
                      raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
                      

                      l_int round-trips correctly, and writing the same pa.Table with pq.write_table
                      preserves the null, so the loss is not pyarrow's.

                      Cause

                      ArrowProjectionVisitor.list rebuilds the array when the element is a struct
                      (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

                      ifisinstance(value_array, pa.StructArray):
                      # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

                      from_arrays receives the offsets buffer alone, which cannot express a null list, and
                      no mask, so the validity bitmap is dropped. That is also why only this one shape is
                      affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
                      does not rebuild at all, and a list whose element is a primitive never enters this
                      branch. The visitor runs on both paths, which is why the same root cause shows up as
                      the read-side assertion above and as the write-side corruption here.

                      Fix

                      Carrying the mask over is enough:

                      list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

                      With that change the reproduction above returns None for both columns, and
                      test_null_list_and_map passes with its commented-out assertion restored. I have not
                      looked at int32-offset or sliced-array handling of this call, which the existing line
                      already relies on; that appears independent of the mask.

                      I have this on a branch with a unit test covering the write path and the integration
                      assertion un-commented, and can open a PR.

                      Found while testing a third-party Iceberg writer against pyiceberg as a reader.

                      Willingness to contribute

                      • I can contribute a fix for this bug independently

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          Null list<struct<...>> is written and read as an empty listΒ #3833

                          Description

                          @slachiewicz

                          Apache Iceberg version

                          main (development)

                          Please describe the bug 🐞

                          A null list<struct<...>> is silently rebuilt as an empty list. This is known on
                          the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
                          asserts the corrupted value, with the correct assertion commented out pending
                          apache/arrow#38809:

                          # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

                          Two things seem worth reporting on top of that.

                          It also affects the write path, where the consequence is worse. The Parquet file
                          pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
                          pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

                          It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
                          takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
                          this particular null loss can be fixed independently of apache/arrow#38809, which is
                          still open.

                          This is the array<struct<>> case from #251. That issue was closed in March 2025 on
                          the strength of this test existing, but the assertion it makes is the corrupted one;
                          the array<int> case in the issue body was genuinely fixed by #252, while the
                          array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
                          remaining broken in
                          #252 (comment) β€” was not.

                          Reproduction (write path)

                          pyiceberg 0.11.1, pyarrow 25.0.1:

                          importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
                          sch=pa.schema([
                          pa.field("id", pa.int32(), nullable=False),
                          pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
                          pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
                          ])
                          tbl=pa.table({"id": [1, 2, 3, 4],
                          "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
                          "l_int": [[1], [], None, [3]]}, schema=sch)
                          cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
                          cat.create_namespace("ns")
                          it=cat.create_table("ns.t", schema=tbl.schema)
                          it.append(tbl)
                          out=it.scan().to_arrow()
                          forcin ("l_struct", "l_int"):
                          print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
                          # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
                          print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

                          Output:

                          l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
                          l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
                          raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
                          

                          l_int round-trips correctly, and writing the same pa.Table with pq.write_table
                          preserves the null, so the loss is not pyarrow's.

                          Cause

                          ArrowProjectionVisitor.list rebuilds the array when the element is a struct
                          (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

                          ifisinstance(value_array, pa.StructArray):
                          # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

                          from_arrays receives the offsets buffer alone, which cannot express a null list, and
                          no mask, so the validity bitmap is dropped. That is also why only this one shape is
                          affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
                          does not rebuild at all, and a list whose element is a primitive never enters this
                          branch. The visitor runs on both paths, which is why the same root cause shows up as
                          the read-side assertion above and as the write-side corruption here.

                          Fix

                          Carrying the mask over is enough:

                          list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

                          With that change the reproduction above returns None for both columns, and
                          test_null_list_and_map passes with its commented-out assertion restored. I have not
                          looked at int32-offset or sliced-array handling of this call, which the existing line
                          already relies on; that appears independent of the mask.

                          I have this on a branch with a unit test covering the write path and the integration
                          assertion un-commented, and can open a PR.

                          Found while testing a third-party Iceberg writer against pyiceberg as a reader.

                          Willingness to contribute

                          • I can contribute a fix for this bug independently

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              Null list<struct<...>> is written and read as an empty listΒ #3833

                              Description

                              @slachiewicz

                              Apache Iceberg version

                              main (development)

                              Please describe the bug 🐞

                              A null list<struct<...>> is silently rebuilt as an empty list. This is known on
                              the read path β€” tests/integration/test_reads.py::test_null_list_and_map currently
                              asserts the corrupted value, with the correct assertion commented out pending
                              apache/arrow#38809:

                              # This should be:# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]# Once https://github.com/apache/arrow/issues/38809 has been fixedassertarrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]

                              Two things seem worth reporting on top of that.

                              It also affects the write path, where the consequence is worse. The Parquet file
                              pyiceberg writes contains an empty list, so the null is gone at rest and no reader β€”
                              pyiceberg, Spark, Trino β€” can recover it. On read the file is at least still correct.

                              It does not depend on the upstream Arrow fix.pa.LargeListArray.from_arrays
                              takes a mask argument β€” since well before pyiceberg's pyarrow>=18.0.0 floor β€” so
                              this particular null loss can be fixed independently of apache/arrow#38809, which is
                              still open.

                              This is the array<struct<>> case from #251. That issue was closed in March 2025 on
                              the strength of this test existing, but the assertion it makes is the corrupted one;
                              the array<int> case in the issue body was genuinely fixed by #252, while the
                              array<struct<test:int>> case in the issue title β€” which @HonahX flagged as
                              remaining broken in
                              #252 (comment) β€” was not.

                              Reproduction (write path)

                              pyiceberg 0.11.1, pyarrow 25.0.1:

                              importos, shutil, globimportpyarrowaspa, pyarrow.parquetaspqfrompyiceberg.catalog.sqlimportSqlCatalogWH="/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
                              sch=pa.schema([
                              pa.field("id", pa.int32(), nullable=False),
                              pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
                              pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
                              ])
                              tbl=pa.table({"id": [1, 2, 3, 4],
                              "l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
                              "l_int": [[1], [], None, [3]]}, schema=sch)
                              cat=SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
                              cat.create_namespace("ns")
                              it=cat.create_table("ns.t", schema=tbl.schema)
                              it.append(tbl)
                              out=it.scan().to_arrow()
                              forcin ("l_struct", "l_int"):
                              print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
                              # the loss is already in the file on disk, not in the read pathf=glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
                              print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())

                              Output:

                              l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
                              l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
                              raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
                              

                              l_int round-trips correctly, and writing the same pa.Table with pq.write_table
                              preserves the null, so the loss is not pyarrow's.

                              Cause

                              ArrowProjectionVisitor.list rebuilds the array when the element is a struct
                              (pyiceberg/io/pyarrow.py:2078 on main @ 7539661):

                              ifisinstance(value_array, pa.StructArray):
                              # This can be removed once this has been fixed:# https://github.com/apache/arrow/issues/38809list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array)

                              from_arrays receives the offsets buffer alone, which cannot express a null list, and
                              no mask, so the validity bitmap is dropped. That is also why only this one shape is
                              affected: the struct visitor passes mask=struct_array.is_null(), the map visitor
                              does not rebuild at all, and a list whose element is a primitive never enters this
                              branch. The visitor runs on both paths, which is why the same root cause shows up as
                              the read-side assertion above and as the write-side corruption here.

                              Fix

                              Carrying the mask over is enough:

                              list_array=pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())

                              With that change the reproduction above returns None for both columns, and
                              test_null_list_and_map passes with its commented-out assertion restored. I have not
                              looked at int32-offset or sliced-array handling of this call, which the existing line
                              already relies on; that appears independent of the mask.

                              I have this on a branch with a unit test covering the write path and the integration
                              assertion un-commented, and can open a PR.

                              Found while testing a third-party Iceberg writer against pyiceberg as a reader.

                              Willingness to contribute

                              • I can contribute a fix for this bug independently

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions