Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Adding introduction, goals, scope and use cases to the RFC - #27

Merged
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases
Sep 13, 2020
Merged

Adding introduction, goals, scope and use cases to the RFC#27
datapythonista merged 8 commits into
masterfrom
purpose_and_use_cases

Conversation

@datapythonista

Copy link
Copy Markdown
Member

This is a first version of the introductory materials of the data frame RFC document.

So far, it only considers the part on data interchange (see #25). The scope will be changed, and further use cases will be added once this first stage in data interchange is completed.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/02_use_cases.md Outdated
Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks @HyukjinKwon for the corrections.

@markusweimermarkusweimer left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just did a pass and left some nits and high-level comments

Comment threadspec/01_purpose_and_scope.md Outdated
[Julia](https://juliadata.github.io/DataFrames.jl/stable/) and others.

In Python, the most popular data frame library is [pandas](https://pandas.pydata.org/).
pandas was initially develop at a hedge fund, with a focus on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

develop -> developed

Comment threadspec/01_purpose_and_scope.md Outdated
implementation details of data frame libraries. This will allow users and third-party libraries to
write code that interacts with a standard data frame, and not with specific implementations.

The defined API does not aim to be a convenient API for all users of data frames. Libraries targeting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we leave the standardization of the end-user API as potential future work for us, or do we not plan on doing any of that?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and one that many readers will have. I think it would be good to explicitly this is out of scope for this version of the standard, but may be in scope for a future version. With a rationale that it's also important, one of the longer-term goals should be (I think) to make the learning curve for users less steep when switching from one library to another one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The structure of:

  • Goals
  • Scope
  • Out-of-scope and non-goals
    is a little inconsistent, I'd suggest to make it symmetric (and add rationales as I just did in my array API scope PR), then this kind of thing may be easier to address.


- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add Database and Big Data systems?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I don't think we are planning to engage with developer of PostgreSQL, MySQL... I'm adding for now big data systems, and also Python libraries to access databases, which I guess we're more likely to engage with. But I'm open to further changes if there are different points of view.

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +124 to +125
Authors of data frame libraries in Python are expected to implement the API defined
in this document in their libraries.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a very heavy handed statement. Could we reword it to something a bit friendlier of:

We encourage data frame libraries in Python to implement the API defined in this document in their libraries

Comment threadspec/01_purpose_and_scope.md Outdated
A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)
- Task schedulers (e.g. Dask, Ray)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would include Mars (https://github.com/mars-project/mars) here as well.

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A dataframe protocol similar to wesm/dataframe-protocol#1 is a prerequisite to this being possible in my mind.

Without having a data exchange protocol defined as part of the spec / goal how can we define from_dataframe / to_dataframe APIs?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point the data exchange protocol is what we're trying to define. This use case tries to illustrate why such a data exchange protocol is needed.

Do you think I should clarify this is the goal for the use cases? Or am I not understanding you?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's good as it is. We are talking about use cases in this document, not the implementation right? So we can loosely define what from_dataframe does, from a high-level point of view, to make the use case clear.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, forgot to comment here. I edited the scope since last comment from @kkraus14. I guess making clear in the goal/scope that we are defining a data exchange protocol solved your concern @kkraus14, or do you think this use case also needs editing?

Thanks both for the comments!

Comment threadspec/02_use_cases.md Outdated
- How data is represented and stored (whether the data is in memory, disk, distributed)
- Expectations on when the execution is happening (in an eager or lazy way)
- Other execution details

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd state here that an API designed for interactive usage is out of scope.

Comment threadspec/01_purpose_and_scope.md Outdated

## Introduction

This document defines a Python data frame API.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer dataframe as one word, like database.

I do not want to start a holy war, and I realize there are historical reasons to call it data frame, but data base was common even throughout the 90s. https://groups.google.com/g/alt.usage.english/c/jRB0g0zK85Q?pli=1

@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback, I addressed all comments.

Comment threadspec/01_purpose_and_scope.md Outdated
- Libraries for database access (e.g. SQLAlchemy)


### Data frame power users

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame power users
### Dataframe power users

Comment threadspec/01_purpose_and_scope.md Outdated
## History
## History and dataframe implementations

Data frame libraries in several programming language exist, such as

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Data frame libraries in several programming language exist, such as
Dataframe libraries in several programming language exist, such as

Comment threadspec/01_purpose_and_scope.md Outdated
This section provides the list of stakeholders considered for the definition of this API.


### Data frame library authors

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
### Data frame library authors
### Dataframe library authors

Comment threadspec/02_use_cases.md
Comment on lines +139 to +144
So, the list above would be reduced to a single function or method in each implementation:

- `from_dataframe()`

Note that the function `from_dataframe()` is for illustration, and not proposed as part
of the standard at this point.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think that it was made clear that a dataframe data exchange protocol was in scope in this document. The only mention of a protocol is in talking about Apache Arrow as far as I can tell.


## Scope

It is in the scope of this document the different elements of the API. This includes signatures

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a verb missing in the first sentence ("to describe" ?)

Comment threadspec/01_purpose_and_scope.md Outdated
Comment on lines +90 to +102
The goal of the API described in this document is to provide a standard interface that encapsulates
implementation details of dataframe libraries. This will allow users and third-party libraries to
write code that interacts with a standard dataframe, and not with specific implementations.

The main goals for the API defined in this document are:

- Provide a common API for dataframes so software can be developed to communicate with it
- Provide a common API for dataframes to build user interfaces on top of it, for example
libraries for interactive use or specific domains and industries
- Simplify interactions between the projects of the ecosystem, for example, software that
receives data as a dataframe
- Make conversion of data among different implementations easier
- Help user transition from one dataframe library to another

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the goal is to initially limit the scope to a "data exchange protocol", the above description of the goals still sound too general, and not making it clear what the specific scope/goals are IMO.

Comment threadspec/01_purpose_and_scope.md Outdated

A non-exhaustive list of upstream categories is next:

- Data formats, protocols and libraries for data analytics (e.g. Apache Arrow)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Numpy as well? It's used by dataframe libraries for their implementation

Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
@maartenbreddels

Copy link
Copy Markdown

I like it!

Co-authored-by: Maarten Breddels <maartenbreddels@gmail.com>
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Comment threadspec/01_purpose_and_scope.md Outdated
Co-authored-by: Devin Petersohn <devin-petersohn@users.noreply.github.com>
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Thanks all for the feedback. If there are no objections I'll be merging this in a couple of days.

I think we will continue iterating on the content of this PR for a bit, particularly the scope. From the last discussions, I think it can be worth adding things like whether we want to support heterogeneous columns, devices, missing values, or virtual columns. But I think it's better to merge this first if people who reviewed it is happy enough, and open follow up PRs. This way the discussion can be more focused, and reviewers don't need to read the long diff of this PR anymore.

@datapythonista
datapythonista merged commit 8352aba into masterSep 13, 2020
@datapythonista
datapythonista deleted the purpose_and_use_cases branch September 13, 2020 21:35
@datapythonista

Copy link
Copy Markdown
MemberAuthor

Merging this. I will keep updating the relevant sections in follow up PRs as needed. Further feedback on what's been merged surely welcome.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@datapythonista@maartenbreddels@rgommers@markusweimer@jorisvandenbossche@kkraus14@HyukjinKwon@devin-petersohn