Happy to move to JIRA if this is confirmed as a bug
In [8]: importpandasaspd
...: importpyarrowasarwIn [9]: df=pd.DataFrame({'A': list('abc'), 'B': np.arange(3)})
...: dfOut[9]:
AB0a01b12c2In [10]: schema=arw.schema([
...: arw.field('A', arw.string()),
...: arw.field('B', arw.int32()),
...: ])
In [11]: tbl=arw.Table.from_pandas(df, preserve_index=False, schema=schema)
...: tblOut[11]:
pyarrow.TableA: stringB: int32metadata--------
{b'pandas': b'{"index_columns": [], "column_indexes": [], "columns": [{"name":'b' "A", "field_name": "A", "pandas_type": "unicode", "numpy_type":'b' "object", "metadata": null}, {"name": "B", "field_name": "B", "'b'pandas_type": "int32", "numpy_type": "int32", "metadata": null}]'b', "pandas_version": "0.23.1"}'}
In [12]: tbl.to_pandas().equals(df)
Out[12]: True...so if the schema matches the pandas datatypes all is well - we can roundtrip the DataFrame.
Now, say we have some bad data such that column 'B' is now of type float64. The datatypes of the DataFrame don't match the explicitly supplied schema object but rather than raising a TypeError the data is silently truncated and the roundtrip DataFrame doesn't match our input DataFame without even a warning raised!
In [13]: df['B'].iloc[0] =1.23
...: dfOut[13]:
AB0a1.231b1.002c2.00In [14]: # I would expect/want this to raise a TypeError since the schema doesn't match the pandas datatypes
...: tbl=arw.Table.from_pandas(df, preserve_index=False, schema=schema)
...: tblOut[14]:
pyarrow.TableA: stringB: int32metadata--------
{b'pandas': b'{"index_columns": [], "column_indexes": [], "columns": [{"name":'b' "A", "field_name": "A", "pandas_type": "unicode", "numpy_type":'b' "object", "metadata": null}, {"name": "B", "field_name": "B", "'b'pandas_type": "int32", "numpy_type": "float64", "metadata": null'b'}], "pandas_version": "0.23.1"}'}
In [15]: tbl.to_pandas() # <-- SILENT TRUNCATION!!!Out[15]:
AB0a11b12c2To be clear, I would really like Table.from_pandas to raise a TypeError if the DataFrame types don't match an explicitly supplied schema and would hope this current behaviour would be considered a bug.
Happy to move to JIRA if this is confirmed as a bug
...so if the
schemamatches the pandas datatypes all is well - we can roundtrip the DataFrame.Now, say we have some bad data such that column 'B' is now of type float64. The datatypes of the DataFrame don't match the explicitly supplied
schemaobject but rather than raising aTypeErrorthe data is silently truncated and the roundtrip DataFrame doesn't match our input DataFame without even a warning raised!To be clear, I would really like
Table.from_pandasto raise aTypeErrorif the DataFrame types don't match an explicitly supplied schema and would hope this current behaviour would be considered a bug.