ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method - #11624

Closed
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608
Closed

ARROW-14608: [Python] Provide access to hash_aggregate functions through a Table.group_by method#11624
amol- wants to merge 26 commits into
apache:masterfrom
amol-:ARROW-14608

Conversation

@amol-

@amol-amol- commented Nov 5, 2021

Copy link
Copy Markdown
Member

No description provided.

@github-actions

Copy link
Copy Markdown

@amol-
amol- marked this pull request as ready for review November 5, 2021 16:26
Comment threadpython/pyarrow/tests/test_table.py Outdated
self._set_options(q, delta, buffer_size, skip_nulls, min_count)


def _group_by(args, keys, aggregations):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can also make this a public function in the compute module?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, should we? I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I made it internal because we plan to replace this with the exec engine on long term, so I guess that the Table.group_by implementation will switch to use something different in the future.

The same could be done for a pyarrow.compute function? (it doesn't map 1:1 to a C++ kernel anyway)

For me one reason to put it in the compute functions as a pc.group_by(table, keys, ...) is to sidestep the 1-step vs 2-step API discussion for the method a bit. For a function in compute, I think it's totally fine to be a one step function

Comment threadpython/pyarrow/tests/test_table.py Outdated
Comment threadpython/pyarrow/tests/test_table.py
Comment threadpython/pyarrow/tests/test_table.py
@ianmcook

Copy link
Copy Markdown
Member

Bikeshedding on the method name: In other packages, the group_by method/function does not actually do any aggregation. Instead it serves as a helper function that tells a separate aggregate method/function what groups to aggregate over. Examples of this include Ibis (group_by --> aggregate), pandas (groupby --> agg), and dplyr (group_by --> summarise). Because of this I think we should pick a different name than group_by for this function, since it both groups and aggregates.

@pitrou

Copy link
Copy Markdown
Member

"grouped_aggregate" perhaps?

@pitrou

Copy link
Copy Markdown
Member

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

Table.group_by would return an intermediate object with several methods, including one for doing the actual grouping ("collect"?) and other(s) to compute aggregates.

@ianmcook

ianmcook commented Nov 12, 2021

Copy link
Copy Markdown
Member

+1 on the two-step approach if it is feasible and doesn't add too much complexity to the implementation.

Ideally the values would be passed to the aggregate function, not to the grouping function. That's how it works in Ibis, dplyr, and pandas (at least since named aggregation in pandas 0.25.0+)

@amol-

Copy link
Copy Markdown
MemberAuthor

Another possibility is to have a two_step API, e.g. replace:

table.group_by("keys", ["values"], "sum")

with:

table.group_by("keys", ["values"]).aggregate("sum")

or perhaps even some shortcuts:

table.group_by("keys", ["values"]).sum()

I think that in such case the aggregated values shouldn't go into group_by, you probably would want something like table.group_by("keys").sum("values").max("othervalues", HashMaxOptions())
I'm not too fond of that solution by the way as it would require an explicit point where you collect results to allow chaining multiple aggregations.

I think having a single aggregate method where you can provide multiple aggregations would be more usable

t.group_by("key").aggregate([
("sum", "values"),
("max", "othervalues", HashMaxOptions())
])

@pitrou

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

(and, yes, you're right, the value columns should go into the aggregate call, not the group_by call. My bad)

@ianmcook

Copy link
Copy Markdown
Member

I was proposing shortcut methods for the simple cases where you compute only one aggregate. But perhaps that's not useful.

Given the small number of aggregate functions and the popularity of that style in pandas, I think that is practical and useful

@jorisvandenbossche

Copy link
Copy Markdown
Member

I am a bit hesitant to add such a two-step interface to pyarrow. It's indeed the way how it is done in other packages, but the ones that @ianmcook mentions (ibis, pandas, dplyr) also all have slightly different APIs on how to specify this. And then pyarrow would add yet another slightly different interface.

(but I also agree that groupby is not a great name as method on the table for this reason)


Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

@amol-

Copy link
Copy Markdown
MemberAuthor

pyarrow would add yet another slightly different interface.
(but I also agree that groupby is not a great name as method on the table for this reason)

I don't have a strong opinion about the single step or multi step API. I personally rarely ever had the need to do a grouping without an associated aggregation, so I feel that the value of the multistep approach isn't huge, even thought it might be easier to evolve in the future.

Playing a bit with this branch, some other observations:

  • I find it unexpected that the resulting table always has "key" column instead of reusing the original name that was specified as the key column
  • Is it possible to group by multiple columns? Not in the current bindings in this PR, but I suppose in c++ / R this is already possible?
  • I think users will very quickly request the ability to specify the resulting column name .. (to not have things like "column_count_distinct")

I implemented support for the first two points in dfecba1
Regarding the third one, I wonder if that would be best satisfied by extending the Table.rename_columns API to support a mapping of column names
IE:

t.rename_column({"oldcolname": "newcolname"})

that might be convenient for other use cases too (for example when willing to rename only a subset of columns) and would expose the ability to do

t.group_by("keycol", ["value1"], ["sum"]).rename_column({"value1_sum": "total"})

@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche@pitrou I was working toward the multistep API refactoring and I was wondering about the grouping alone case (group_by(["keys"]).collect()?).

At the moment it seems that the GroupBy C++ function doesn't support grouping without any provided aggregation. Do you think it would be a reasonable work-around to run a count aggregation just to drop it or should we just leave out the plain grouping for the moment?

@pitrou

Copy link
Copy Markdown
Member

IMHO we should just leave out the plain grouping for the moment.

Comment threadpython/pyarrow/table.pxi Outdated
@amol-

amol- commented Nov 17, 2021

Copy link
Copy Markdown
MemberAuthor

@pitrou moved to multistep api

 table.group_by("keys").aggregate([
("sum", "values"),
("count", "values")
])

or

 table.group_by("keys").aggregate([
("sum", "values", FunctionOptions),
("count", "values", FunctionOptions)
])

at the moment it only has aggregate method, but we can grow more helpers in the future

@chunggchungg left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure what other functionality we plan on supporting but syntax makes sense to me. it's similar to other dataframe libraries.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
@alippai

Copy link
Copy Markdown
Contributor

A slightly related silly question: how does the performance compare to pandas at this stage?

@amol-

Copy link
Copy Markdown
MemberAuthor

A slightly related silly question: how does the performance compare to pandas at this stage?

Not a real benchmark, but in the current very rough form it seems to be mostly comparable.

>>> table = pyarrow.csv.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: table.group_by("VendorID").aggregate([("sum", "trip_distance")]), number=1)
2.626802896999993

VS

>>> df = pandas.read_csv("yellowtaxi.csv")
>>> timeit.timeit(lambda: df.groupby("VendorID").aggregate({"trip_distance": ["sum"]}), number=1)
2.3642018030000003

@pitroupitrou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a nit. @jorisvandenbossche Can you give this a final review?

function_registry,
get_function,
list_functions,
_group_by

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reaons for exposing this publicly? Is this just a leftover from previous attempts?

@amol-amol-Nov 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's to make it available when _pc() is used to access compute functions from other modules/files (in this case from table.pxi). I adhered to that practice instead of injecting a import pyarrow._compute.

Given that there are many more internal functions in the pyarrow.compute module I thought it wasn't concerning.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi Outdated
amol-and others added 3 commits November 19, 2021 10:25
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, a few more (mostly docstring) nits.

Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/table.pxi
Comment threadpython/pyarrow/table.pxi Outdated
Comment threadpython/pyarrow/tests/test_table.py
amol-and others added 2 commits November 23, 2021 12:39
Co-authored-by: Joris Van den Bossche <jorisvandenbossche@gmail.com>
@amol-

Copy link
Copy Markdown
MemberAuthor

@jorisvandenbossche I should have addressed your most recent comments, anything else you feel is pending?

@jorisvandenbosschejorisvandenbossche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

@ursabot

ursabot commented Nov 25, 2021

Copy link
Copy Markdown

Benchmark runs are scheduled for baseline = 4cfd1d9 and contender = 999d97a. 999d97a is a master commit associated with this PR. Results will be available as each benchmark for each run completes.
Conbench compare runs links:
[Finished ⬇️0.0% ⬆️0.0%] ec2-t3-xlarge-us-east-2
[Failed] ursa-i9-9960x
[Finished ⬇️0.18% ⬆️0.0%] ursa-thinkcentre-m75q
Supported benchmarks:
ursa-i9-9960x: langs = Python, R, JavaScript
ursa-thinkcentre-m75q: langs = C++, Java
ec2-t3-xlarge-us-east-2: cloud = True

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@amol-@ianmcook@pitrou@jorisvandenbossche@alippai@ursabot@lidavidm@nealrichardson@chungg