Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

spin

Lifecycle: experimental

spin – simulating prediction intervals – is a toy package for building prediction intervals with out-of-sample data for tidymodels supported workflows objects. It consists of two functions: prep_interval() for simulating model fits + sample uncertainty and predict_interval() for outputting prediction intervals on new observations.

To use spin, you need a dataset for training and a tidymodels workflow (with a preprocessing recipe and model specification):

# Set-up
library(tidyverse)
library(tidymodels)
library(spin)
library(earth)
set.seed(123)
split<-palmerpenguins::penguins %>% na.omit() %>% initial_split()
train<- training(split)
other_data<- testing(split)
# Specify recipe and model and combine into workflowrec<- recipe(body_mass_g~., data=train) %>% step_scale(all_numeric_predictors())
mod<-parsnip::mars(prod_degree=3) %>% set_engine("earth") %>% set_mode("regression")
workflow<-workflows::workflow() %>% add_recipe(rec) %>% add_model(mod) 

prep_interval() takes in the workflow and a dataset that will be used for estimating prediction intervals based on out-of-sample errors.

set.seed(12)
# output for a 95% prediction intervalintervals_prepped<-workflow %>% prep_interval(fit_data=train)

The resulting object is passed into predict_interval() which can output prediction intervals on new_data.

set.seed(12)
pred_intervals<-intervals_prepped %>%
predict_interval(new_data=other_data,
probs= c(0.025, 0.50, 0.975))
pred_intervals#> # A tibble: 83 x 3#> probs_0.025 probs_0.500 probs_0.975#> <dbl> <dbl> <dbl>#> 1 3153. 3858. 4492.#> 2 2672. 3414. 4070.#> 3 3361. 4234. 4971.#> 4 2702. 3402. 4110.#> 5 3347. 4177. 5520.#> 6 2397. 3197. 3921.#> 7 3271. 3970. 4589.#> 8 2745. 3489. 4138.#> 9 3191. 3954. 4587.#> 10 2727. 3392. 4026.#> # ... with 73 more rows

Let’s overlay the actual body_mass_g on the prediction intervals:

bind_cols(
select(other_data, actuals=body_mass_g),
set_names(pred_intervals,
paste0(".pred", c("_lower", "", "_upper")))
) %>% ggplot(aes(x=.pred, y=actuals))+
geom_point(aes(y=.pred, color="prediction interval"))+
geom_errorbar(aes(ymin=.pred_lower, ymax=.pred_upper, color="prediction interval"))+
geom_point(aes(y=actuals, color="actuals"))+
theme_bw()+
labs(title="Simulated Prediction Intervals for body_mass_g",
subtitle="For Palmer's Penguins",
y="body_mass_g")

Intervals seem to have pretty good coverage, even on this holdout data.

When to use

Building robust prediction intervals usually happens after model selection, hence the workflow passed into predict_interval() should have all hyperparameters etc. specified[1].

More Documentation

For a more detailed description and review of the current methodology used in spin, see Simulating Prediction Intervals.

Notes, Limitations, Resources, Ideas

spin produces reasonable, simple to create prediction interval (assuming iid) for any model type supported by the tidymodels ecosystem. Due to it taking in a workflow, spin takes an expansive view of uncertainty due to model estimation and include both model fit as well as preprocessing steps (which feels like a good idea in most cases). That said, spin is very much in toy / experimental form[2]:

  • A more detailed description of the (rough) methodology currently used by spin can be found in the second of three blog posts I wrote on building prediction intervals: Simulating Prediction Intervals where I also include references to related posts by Dan Saattrup Nielsen and others (Dan has since published the python package Doubt). See parsnip#464 for other related links.
  • The methodology in spin assumes errors are roughly iid[3]. See other post on quantile regression intervals and linked to resources for other (in many circumstances) more flexible approaches.
  • I’d like to investigate blending a parsnip interface like spin onto a conformal inference approach
  • I’d like to allow for more flexibility in specifying things like the resampling methodology (i.e. if have custom resampling scheme that makes since for the problem, e.g. in time series), allowing errors to be non-iid, …
  • Default builds sqrt(nrow(data)) models – which can take a long time depending on number of observations and model type. May want to set-up for doing in parallel or look into other approaches for speed-ups.

Installation

# install.packages("devtools")devtools::install_github("brshallo/spin")

[1] Though spin may also be useful for evaluating prediction intervals of different candidate models.

[2] There are also no checks, tests, etc currently in place.

[3] For example, heteroskedasticity, like that shown in this gist does not produce good interval fits.

About

Package for Simulating Prediction Intervals

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages