Skip to content

Add Ellipsis typehint to reductions - #7048

Merged
dcherian merged 12 commits into
pydata:mainfrom
headtr1ck:ellipsis
Sep 28, 2022
Merged

Add Ellipsis typehint to reductions#7048
dcherian merged 12 commits into
pydata:mainfrom
headtr1ck:ellipsis

Conversation

@headtr1ck

Copy link
Copy Markdown
Collaborator

This PR adds the ellipsis typehint to reductions (only where they behave differently from None to reduce overhead).
Follow up on #7017 (comment)

Additionally I was changing a lot of "one or more dimensions" typehints to str | Iterable[Hashable] (See #6142).
Some code changes were necessary to support this fully. Before several things were not working with actual hashable dimensions that are not strings.

@headtr1ckheadtr1ck changed the title EllipsisAdd Ellipsis typehint to reductionsSep 16, 2022
@max-sixty

Copy link
Copy Markdown
Collaborator

Excellent @headtr1ck !

Do we need to run pytest --accept to get the docstrings? It looks like we lost lots...

@headtr1ck

Copy link
Copy Markdown
CollaboratorAuthor

Could any dev that uses linux rerun the generate_reductions and pytest --doctest-modules xarray/core/_reductions.py --accept?
On windows I still get different results (maybe that should be fixed at some point...)

Comment threadxarray/util/generate_reductions.py Outdated
@headtr1ck

Copy link
Copy Markdown
CollaboratorAuthor

Turns out that the buildin ellipsis works now with mypy.
Did not test it for older python versions, may require some special casing (Lets see if the tests pass)?

@max-sixty

Copy link
Copy Markdown
Collaborator

Here's the diff from pytest-accept (it is weird that it's slightly different on windows...)

commit 83615a94a6b7c0ae0cf0e0240d7705d9ce6c21e5
Author: Maximilian Roos <m@maxroos.com>
Date: Sat Sep 17 12:36:45 2022 -0700
pytest-accept
diff --git a/xarray/core/_reductions.py b/xarray/core/_reductions.py
index a7cf7ec2..d0c2a9d7 100644
--- a/xarray/core/_reductions.py+++ b/xarray/core/_reductions.py@@ -97,7 +97,7 @@ def count(
<xarray.Dataset>
Dimensions: ()
Data variables:
- da int32 5+ da int64 5
"""
return self.reduce(
duck_array_ops.count,
@@ -4400,7 +4400,7 @@ def count(
>>> da.groupby("labels").count()
<xarray.DataArray (labels: 3)>
- array([1, 2, 2], dtype=int64)+ array([1, 2, 2])
Coordinates:
* labels (labels) object 'a' 'b' 'c'
"""
@@ -5485,7 +5485,7 @@ def count(
>>> da.resample(time="3M").count()
<xarray.DataArray (time: 3)>
- array([1, 3, 1], dtype=int64)+ array([1, 3, 1])
Coordinates:
* time (time) datetime64[ns] 2001-01-31 2001-04-30 2001-07-31
"""

@IllviljanIllviljan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if Dims should include ellipsis as well? The few times it's missing might be issues with the functions?

Comment threadxarray/core/dataarray.py Outdated
Comment threadxarray/core/dataarray.py Outdated
Comment threadxarray/core/dataarray.py Outdated
Comment threadxarray/core/dataarray.py Outdated
Comment threadxarray/core/dataset.py Outdated
Comment threadxarray/core/variable.py Outdated
Comment threadxarray/core/variable.py Outdated
Comment threadxarray/core/variable.py Outdated
Comment threadxarray/core/variable.py Outdated
Comment threadxarray/core/weighted.py Outdated
@Illviljan

Illviljan commented Sep 18, 2022

Copy link
Copy Markdown
Contributor
@@ -4400,7 +4400,7 @@ def count(
>>> da.groupby("labels").count()
<xarray.DataArray (labels: 3)>
- array([1, 2, 2], dtype=int64)+ array([1, 2, 2])
Coordinates:
* labels (labels) object 'a' 'b' 'c'
"""

Is it just me that this example crashes the second time I run it?

importnumpyasnpimportpandasaspdimportxarrayasxrda=xr.DataArray(
np.array([1, 2, 3, 1, 2, np.nan]),
dims="time",
coords=dict(
time=("time", pd.date_range("01-01-2001", freq="M", periods=6)),
labels=("time", np.array(["a", "b", "c", "c", "b", "a"])),
),
)
da.groupby("labels").count()
Details
Traceback (mostrecentcalllast):
File"C:\Users\J.W\anaconda3\envs\xarray-tests\lib\site-packages\spyder_kernels\py3compat.py", line356, incompat_execexec(code, globals, locals)
File"g:\program\dropbox\python\xarray_groupby_windows_diff.py", line34, in<module>da.groupby("labels").count()
File"c:\users\j.w\documents\github\xarray\xarray\core\_reductions.py", line4384, incountreturnself._flox_reduce(
File"c:\users\j.w\documents\github\xarray\xarray\core\groupby.py", line738, in_flox_reduceresult=xarray_reduce(
File"C:\Users\J.W\anaconda3\envs\xarray-tests\lib\site-packages\flox\xarray.py", line240, inxarray_reduceds, *by=xr.broadcast(ds, *by, exclude=exclude_dims)
File"c:\users\j.w\documents\github\xarray\xarray\core\alignment.py", line1046, inbroadcastargs=align(*args, join="outer", copy=False, exclude=exclude)
File"c:\users\j.w\documents\github\xarray\xarray\core\alignment.py", line765, inalignaligner.align()
File"c:\users\j.w\documents\github\xarray\xarray\core\alignment.py", line549, inalignself.find_matching_indexes()
File"c:\users\j.w\documents\github\xarray\xarray\core\alignment.py", line256, infind_matching_indexesobj_indexes, obj_index_vars=self._normalize_indexes(obj.xindexes)
File"c:\users\j.w\documents\github\xarray\xarray\core\alignment.py", line205, in_normalize_indexespd_idx=safe_cast_to_index(data)
File"c:\users\j.w\documents\github\xarray\xarray\core\utils.py", line140, insafe_cast_to_indexindex=pd.Index(np.asarray(array), **kwargs)
File"C:\Users\J.W\anaconda3\envs\xarray-tests\lib\site-packages\pandas\core\indexes\base.py", line483, in__new__data=sanitize_array(data, None, dtype=dtype, copy=copy)
File"C:\Users\J.W\anaconda3\envs\xarray-tests\lib\site-packages\pandas\core\construction.py", line524, insanitize_arrayraiseValueError("index must be specified when data is not list-like")
ValueError: indexmustbespecifiedwhendataisnotlist-like

@headtr1ck

Copy link
Copy Markdown
CollaboratorAuthor

Is it just me that this example crashes the second time I run it?

Could you specify what you mean by "second time I run it"?
Executing the groupby twice?

@Illviljan

Copy link
Copy Markdown
Contributor

Just running that script file several times without restarting the console. It might be a Spyder bug though since I can't reproduce it in a stand alone ipython console.

This for example works (the first time):

importnumpyasnpimportpandasaspdimportxarrayasxrda=xr.DataArray(
np.array([1, 2, 3, 1, 2, np.nan]),
dims="time",
coords=dict(
time=("time", pd.date_range("01-01-2001", freq="M", periods=6)),
labels=("time", np.array(["a", "b", "c", "c", "b", "a"])),
),
)
da.groupby("labels").count()
da=xr.DataArray(
np.array([1, 2, 3, 1, 2, np.nan]),
dims="time",
coords=dict(
time=("time", pd.date_range("01-01-2001", freq="M", periods=6)),
labels=("time", np.array(["a", "b", "c", "c", "b", "a"])),
),
)
da.groupby("labels").count()

Comment threadxarray/core/dataset.py
@headtr1ck

Copy link
Copy Markdown
CollaboratorAuthor
importnumpyasnpimportpandasaspdimportxarrayasxrda=xr.DataArray(
np.array([1, 2, 3, 1, 2, np.nan]),
dims="time",
coords=dict(
time=("time", pd.date_range("01-01-2001", freq="M", periods=6)),
labels=("time", np.array(["a", "b", "c", "c", "b", "a"])),
),
)
da.groupby("labels").count()
da=xr.DataArray(
np.array([1, 2, 3, 1, 2, np.nan]),
dims="time",
coords=dict(
time=("time", pd.date_range("01-01-2001", freq="M", periods=6)),
labels=("time", np.array(["a", "b", "c", "c", "b", "a"])),
),
)
da.groupby("labels").count()

For me this works in a python terminal, python script and jupyter notebook (I don't use spyder but vscode).

@headtr1ckheadtr1ck added the plan to merge Final call for comments label Sep 25, 2022
@headtr1ckheadtr1ck mentioned this pull request Sep 27, 2022
3 tasks
@dcherian
dcherian merged commit 226c23b into pydata:mainSep 28, 2022
@headtr1ck
headtr1ck deleted the ellipsis branch September 28, 2022 18:02
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@headtr1ck@max-sixty@Illviljan@dcherian