Skip to content

[SPARK-11057] [SQL] Add correlation and covariance matrices - #9366

Closed
NarineK wants to merge 1 commit into
apache:masterfrom
NarineK:sparksqlcorcov
Closed

[SPARK-11057] [SQL] Add correlation and covariance matrices#9366
NarineK wants to merge 1 commit into
apache:masterfrom
NarineK:sparksqlcorcov

Conversation

@NarineK

Copy link
Copy Markdown
Contributor

Hi there,

As we know R has the option to calculate the correlation and covariance for all columns of a dataframe or between columns of two dataframes.

If we look at apache math package we can see that, they have that too.
http://commons.apache.org/proper/commons-math/apidocs/org/apache/commons/math3/stat/correlation/PearsonsCorrelation.html#computeCorrelationMatrix%28org.apache.commons.math3.linear.RealMatrix%29

In case we have as input only one DataFrame:


for correlation:
cor[i,j] = cor[j,i]
and for the main diagonal we can have 1s.


for covariance:
cov[i,j] = cov[j,i]
and for main diagonal: we can compute the variance for that specific column:
See:
http://commons.apache.org/proper/commons-math/apidocs/org/apache/commons/math3/stat/correlation/Covariance.html#computeCovarianceMatrix%28org.apache.commons.math3.linear.RealMatrix%29

Thanks,
Narine

@NarineK

Copy link
Copy Markdown
ContributorAuthor

@shivaram , @rxin , would you guys, please, take a look at this ?
Thanks!

@shivaram

Copy link
Copy Markdown
Contributor

cc @mengxr

@SparkQA

Copy link
Copy Markdown

Test build #44651 has finished for PR 9366 at commit 74bdf54.

  • This patch passes all tests.
  • This patch merges cleanly.
  • This patch adds no public classes.

@NarineK

Copy link
Copy Markdown
ContributorAuthor

Hi guys, would you share your thoughts about this ?
Thanks!

@NarineK

Copy link
Copy Markdown
ContributorAuthor

In general I think that currently there are some issues in the StatFunctions.scala:

It seems that all computations both for covariance and correlation are being accomplished in one place which makes it a little confusing and harder to extend for the future.

collectStatisticalData method is called for both correlation and covariance and even if I call something like this:
df.stats.corr("numeric_colame", "string_colname")
I get an error like this:
java.lang.IllegalArgumentException: requirement failed: Covariance calculation for columns with dataType StringType not supported.

Here is an example:
These 2 variables are being computed each time when we compute covariance, however, are being used only for correlation:
var MkX = 0.0 // sum of squares of differences from the (current) mean for col1
var MkY = 0.0 // sum of squares of differences from the (current) mean for col2

I think we can actually separate the computations. Is there a reason why these computations are being accomplished in one place ? @rxin, @mengxr

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can't assume all columns are of numeric type. Catch exception here and use null as value if exception happens?

@NarineK

Copy link
Copy Markdown
ContributorAuthor

Hi @sun-rui,
thank you for your comment. In general, I think that, it might be better to verify all columns types and make sure that we are dealing with numeric fields. if any of the fields isn't numeric we can show an error message, similar to R.
cor(iris)
Error in cor(iris) : 'x' must be numeric

@NarineK

Copy link
Copy Markdown
ContributorAuthor

what do you think ?

@sun-rui

Copy link
Copy Markdown
Contributor

Yes, since R throws error message in this case, we can leave exception un-handled. No need to verify all column types. User will get exception message at https://github.com/apache/spark/blob/master/sql/core/src/main/scala/org/apache/spark/sql/execution/stat/StatFunctions.scala#L81

@NarineK

Copy link
Copy Markdown
ContributorAuthor

yes, there is even a test case which covers that case.

@NarineK

Copy link
Copy Markdown
ContributorAuthor

can someone from Spark SQL committers or experts also look at this ?

@SparkQA

Copy link
Copy Markdown

Test build #53308 has finished for PR 9366 at commit 74bdf54.

  • This patch fails R style tests.
  • This patch does not merge cleanly.
  • This patch adds no public classes.

@SparkQA

Copy link
Copy Markdown

Test build #56142 has finished for PR 9366 at commit 74bdf54.

  • This patch fails R style tests.
  • This patch does not merge cleanly.
  • This patch adds no public classes.

@shivaram

Copy link
Copy Markdown
Contributor

cc @mengxr

@sjjpo2002

sjjpo2002 commented Apr 22, 2016

Copy link
Copy Markdown

I have been trying to use correlation on a matrix with many columns. @NarineK menthioned R like correlation. I wish we had something like what pandas offers. It handles missing data automatically. Take a look here. Even the corr() function from MLlib can not handle missing data. These features are really missing from SparkSQL:

  • Apply correlation on all columns and return a matrix
  • Handle missing data automatically like how pandas does

@gatorsmile

Copy link
Copy Markdown
Member

@NarineK Are you still working on this? cc @yanboliang

@HyukjinKwonHyukjinKwon mentioned this pull request Jun 25, 2017
@gatorsmile

Copy link
Copy Markdown
Member

We are closing it due to inactivity. please do reopen if you want to push it forward. Thanks!

zifeif2 pushed a commit to zifeif2/spark that referenced this pull request Nov 22, 2025
## What changes were proposed in this pull request?
This PR proposes to close stale PRs, mostly the same instances with apache#18017
I believe the author in apache#14807 removed his account.
Closesapache#7075Closesapache#8927Closesapache#9202Closesapache#9366Closesapache#10861Closesapache#11420Closesapache#12356Closesapache#13028Closesapache#13506Closesapache#14191Closesapache#14198Closesapache#14330Closesapache#14807Closesapache#15839Closesapache#16225Closesapache#16685Closesapache#16692Closesapache#16995Closesapache#17181Closesapache#17211Closesapache#17235Closesapache#17237Closesapache#17248Closesapache#17341Closesapache#17708Closesapache#17716Closesapache#17721Closesapache#17937
Added:
Closesapache#14739Closesapache#17139Closesapache#17445Closesapache#18042Closesapache#18359
Added:
Closesapache#16450Closesapache#16525Closesapache#17738
Added:
Closesapache#16458Closesapache#16508Closesapache#17714
Added:
Closesapache#17830Closesapache#14742
## How was this patch tested?
N/A
Author: hyukjinkwon <gurwls223@gmail.com>
Closesapache#18417 from HyukjinKwon/close-stale-pr.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@NarineK@shivaram@SparkQA@sun-rui@sjjpo2002@gatorsmile