Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

R-Package Build Status

News: Version 2.0

  • Improved support for multinomial logistic and Cox proportional hazards regression.
  • Improved performance (runtime avx detection and multithreading support now also available for macOS).
  • The function zeroSumCVFit() and zeroSumFit() have been merged to zeroSum().
  • The function arguments have been renamed to correspond to the glmnet R package. Please see the description below and the examples shown in the corresponding example section.

Introduction: The R-Package zeroSum

zeroSum is an R-package for fitting scale invariant and thereby reference point insensitive log-linear models by imposing the zero-sum constraint [1] combined with the elastic-net regularization [4].

The zero-sum constraint is recommended for fitting linear models between a response yi and log-transformed data xi where ambiguities in the reference point translate to sample-wise shifts. The influence of such sample-wise shifts γi on linear models is as follows:

By restricting the sum of coefficients (red) to zero the model becomes reference point insensitive.

This approach of has been proposed in the context of compositional data in [3] and in the context of reference points in [1]. The corresponding minimization problem reads:

where L denotes the log-likelihood function of the regression type. The parameter α can be used to adjust the ratio between ridge and LASSO regularization. For α=0 the elastic-net becomes a ridge regularization, for α=1 the elastic-net becomes the LASSO regularization.

For more details about zero-sum see [1] and [2].

The function calls of the zeroSum package follow closely those of the glmnet package [4,5,6]. Therefore, the results of the linear models with or without the zero-sum constraint can be easily compared.

Please note that the zero-sum constraint only yields reference point insensitive models on log transformed data!

Table of contents

Installation

Dependencies

  • The R-package Matrix is required for sparse matrices. (Part of R recommended and should be installed by default)
  • The R-package testthat is only required for unit testing.
  • For Installation from source or for using devtools::install_github() a modern compiler which supports C++11 and AVX512 intrinsics is required.

Windows

A binary package (*.zip) is available, which can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.zip", repos = NULL)

Linux and macOS

The source package (*.tar.gz) can be installed in R using:

install.packages("https://github.com/rehbergT/zeroSum/raw/main/zeroSum_2.0.7.tar.gz", repos = NULL)

Devtools

The devtools package allows to install packages from github:

devtools::install_github("rehbergT/zeroSum/zeroSum")

Manually building from source

Git clone or download this repository and open a terminal within the zeroSum folder. The source package can be build and installed with:

R CMD build zeroSum
R CMD INSTALL zeroSum_2.0.7.tar.gz

Quick start

Load the zeroSum package and the included example data

library(zeroSum) # load the R package
set.seed(3) # set a seed for exact same results
x <- log2(exampleData$x) # load example data and use log transformation
y <- exampleData$y

Perform a cross-validation over an automatically approximated λ sequence to determine the optimal value of the elastic-net regularization:

fit <- zeroSum(x, y)

use the plot() function to see the CV-error versus the regularization strength λ:

plot(fit)

The lowest mean squared error indicates the best choice for λ.

The coef() function can be used to extract the coefficients:

coef(fit)
> 49 x 1 sparse Matrix of class "dgCMatrix"
>
> intercept 37.22608690
> Feature1 -0.32431232
> Feature2 1.00806569
> Feature3 .
> Feature4 .
> Feature5 1.35545432
> ...

Note that the sum of coefficients is zero:

sum(coef(fit)[-1,]) # -1 to remove the intercept
> 8.326673e-17

Use the predict() function to make predictions:

predict(fit, newx=x)
> [,1]
> Sample1 -5.502157
> Sample2 47.787166
> Sample3 31.794299
> Sample4 36.524800
> Sample5 41.494367
> ...

Supported regression types

  • Linear regression (family = "gaussian", default):

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, family = "gaussian")
    
  • Binomial logistic regression (family = "binomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    fit <- zeroSum(x, y, family = "binomial")
    
  • Multinomial logistic regression (family = "multinomial")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yMultinomial # y is a integer vector containing 1, 2, 3, ...
    # representing the different classes
    fit <- zeroSum(x, y, family = "multinomial")
    
  • Cox proportional hazard regression (family = "cox")

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$yCox # y is a numeric matrix with two columns and N
    # rows. The first column depicts the observation
    # time and the second column indicates death (1)
    # or right censoring (0).
    fit <- zeroSum(x, y, family = "cox")
    

FAQ

  • Adjusting the elastic-net parameter α:

    The parameter α is default set to 1 (LASSO case) and can be adjused with the argument alpha. For example:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, alpha = 0.9)
    
  • Manually defining the folds of the cross-validation in order to get reproducible results:

    The argument foldid can be used to manually define the folds of the internal cross-validation which is used to determine the optimal value of λ. Thereby, repeated runs yield the exact same result

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed to get the same foldids
    foldid <- sample(rep(1:10, length.out = nrow(x))) # randomly assign a fold id for each sample
    # (10 folds in total)
    fit1 <- zeroSum(x, y, foldid = foldid)
    fit2 <- zeroSum(x, y, foldid = foldid)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 32.8706621
    > Feature1 . .
    > Feature2 1.2074393 1.2074393
    > Feature3 . .
    > Feature4 0.1406724 0.1406724
    > Feature5 0.9506165 0.9506165
    > Feature6 0.5565038 0.5565038
    > ...
    

    Without pre-defining the folds each zeroSum call will automatically generate random fold ids causing that repeated runs can yield different results.

     fit3 <- zeroSum(x, y)
    fit4 <- zeroSum(x, y)
    cbind(coef(fit3), coef(fit4))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.49787458 32.005580496
    > Feature1 -0.06557239 -0.109776390
    > Feature2 1.23434942 1.227301341
    > Feature3 . .
    > Feature4 0.34301075 1.085574015
    > Feature5 0.85595893 0.586373614
    > Feature6 0.45418722 .
    
  • Using sample-wise weights

    The argument weights allows to adjust the contribution of each sample to the log-likelihood. This allows for instance to balance different group sizes:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$ylogistic # y is a integer vector containing 1 and 0
    set.seed(1) # set a seed to get the same foldids
    c(sum(y == 0), sum(y == 1)) # the samples are almost balanced
    > 16 15
    weights <- rep(0, nrow(x)) # generate a weight vector where the length
    # corresponds to the number of samples
    weights[y == 0] <- rep(1/16, 16)
    weights[y == 1] <- rep(1/15, 15)
    fit <- zeroSum(x, y, weights=weights, family = "binomial")
    
  • Excluding features from the elastic-net regularization using penalty.factor

    The argument penalty.factor (vi) allows to adjust the contribution of a feature to the elastic-net regualization: (By default each factor is 1)

    For instance, setting the first value of the penalty.factor to 0 causes that the first coefficient is not affected by the regularization and will therefore be non-zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    pf <- rep(1, ncol(x))
    pf[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, penalty.factor = pf)
    cbind(coef(fit1), coef(fit2))
    > 49 x 2 sparse Matrix of class "dgCMatrix"
    >
    > intercept 32.8706621 31.83764044
    > Feature1 . -2.80130699
    > Feature2 1.2074393 0.91933940
    > Feature3 . .
    > Feature4 0.1406724 1.24848083
    > Feature5 0.9506165 0.64844358
    > Feature6 0.5565038 .
    > ...
    

    A penalty.factor of 0 for a specific feature in combination with a zero-sum weight of 0 (see next point) for the same feature can be used for confounding features.

  • Excluding features from the zero-sum constrain using zeroSum.weights

    The argument zeroSum.weights (ui) allows to adjust the zero-sum constraint. Each factor of the constraint is multiplied with with corresponding coefficient: (By default each factor is 1)

    For instance, setting the first value of the zeroSum.weights to 0 causes that the first coefficient is not affected by the zero-sum constraint. Thus, the sum of coefficents excluding the first coefficient will be zero:

     library(zeroSum) # load the R package
    x <- log2(exampleData$x) # load example data and apply log2()
    y <- exampleData$y # y is a numeric vector
    set.seed(1) # set a seed
    zw <- rep(1, ncol(x))
    zw[1] <- 0
    fit1 <- zeroSum(x, y)
    fit2 <- zeroSum(x, y, zeroSum.weights = zw)
    sum(coef(fit1)[-1, ]) # -1 because of the intercept
    > [1] -3.996803e-15
    sum(coef(fit2)[-1, ]) # -1 because of the intercept
    [1] -0.1745896 # non zero because the first element
    # is excluded from the constraint
    sum(coef(fit2)[-c(1:2), ]) # the sum of all features except
    > [1] -3.996803e-15 # the intercept and the first feature is 0
    

    A zeroSum.weight of 0 for a specific feature in combination with a penalty.factor of 0 for the same feature can be used for confounding features.

  • Disabling the zero-sum contraint:

    The zero-sum constraint can be disabled by using the argument zeroSum = FALSE, which should yield results equivalent to the glmnet package with standardize = FALSE. For example:

     library(zeroSum) # load the R package
    set.seed(1)
    x <- log2(exampleData$x) # load example data and use log transformation
    y <- exampleData$y # y is a numeric vector
    fit <- zeroSum(x, y, zeroSum = FALSE)
    

    The sum of all coefficients (without the intercept) is thus non-zero:

     sum(coef(fit)[-1, ])
    > [1] -0.7952836
    

Standalone HPC version

zeroSum is available as a standalone C++ program which allows an easy utilization on computing clusters. Note that the mpi parallelization can only be used for the fused LASSO regularization.

Build instructions

Git clone or download this repository and a open a terminal within the zeroSum/hpc_version folder. Build the zeroSum hpc version with:

make

By default the Makefile is configured to compile zeroSum for the currently used architecture (native). This can be adjusted at the top of the Makefile.

Basic usage

We provide R functions for exporting all necessary files for the standalone HPC version and for importing the generated results.

Open a terminal and create an empty folder. Change the working directory to the created folder. Load zeroSum and the example data:

library(zeroSum)
set.seed(1)
x <- log2(exampleData$x)
y <- exampleData$y

The exportToCSV() function has the same arguments as zeroSum() but with the additional arguments name and path. With name one can set the prefix name of all exported files and with path the path for exporting the files can be determined.

exportToCSV( x, y, name="test", path="./" )

4 files are now generated:

$ ls
rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

The .rds file is only necessary for importing the results and only the other 3 have to be passed to the HPC zeroSum version as follows:

mpirun -np 1 /path/to/hpcVersion/zeroSum settingstest.csv xtest.csv ytest.csv ./ example

The last two arguments determine the path where the results should be saved and the of name the file in which the results should be saved:

$ ls
example_0_stats.csv rDataObject_test.rds settingstest.csv xtest.csv ytest.csv

As one can see the file example_0_stats.csv containing the results has been created. One can now import the file back into R with the importFromCSV() function:

library(zeroSum)
cv.fit <- importFromCSV("rDataObject_test.rds", "example" )
head( coef(cv.fit, s="lambda.min") )
[,1]
intercept 27.3026110
Feature1 0.0000000
Feature2 0.7558415
Feature3 0.0000000
Feature4 0.3207761
Feature5 0.5220436

The same coefficients as above have been obtained (despite some numerical uncertainty).

References

[1] M. Altenbuchinger, T. Rehberg, H. U. Zacharias, F. Staemmler, K. Dettmer, D. Weber, A. Hiergeist, A. Gessner, E. Holler, P. J. Oefner, R. Spang. Reference point insensitive molecular data analysis. Bioinformatics 33(2):219, 2017. doi: 10.1093/bioinformatics/btw598

[2] H. U. Zacharias, T. Rehberg, S. Mehrl, D. Richtmann, T. Wettig, P. J. Oefner, R. Spang, W. Gronwald, M. Altenbuchinger. Scale-invariant biomarker discovery in urine and plasma metabolite fingerprints. J. Proteome Res. doi: 10.1021/acs.jproteome.7b00325, ArXiv e-prints. http://arxiv.org/abs/1703.07724

[3] Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 2014. doi: 10.1093/biomet/asu031.

[4] H. Zou and T. Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301-320, 2005.

[5] J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1-22, 2010. ISSN 1548-7660. doi: 10.18637/jss.v033.i01.

[6] N. Simon, J. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software, 39(1):1-13, 2011. ISSN 1548-7660. doi: 10.18637/jss.v039.i05.

About

R-package for elastic net regularized regression with zero sum constraint

Resources

Stars

27 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages