Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

lightMLFlow

A lightweight R wrapper for the MLFlow REST API

R-CMD-checkLifecycle: experimental

Setup

lightMLFlow will soon be on CRAN!

For now, you can install it from Github as follows:

# install.packages("devtools")devtools::install_github("collegevine/lightMLFlow")

The Package

This package differs from the CRAN mlflow package in a few important ways.

First, there are some things that the full CRAN mlflow package has that this package doesn’t:

  • mlflow supports more configurable artifact stores, while lightMLFlow only supports S3 as the artifact store. That said, we would love contributions from Azure or GCP users to extend this functionality!
  • mlflow allows you to run install_mlflow() to install MLFlow on your machine. lightMLFlow assumes you’re using an MLFlow instance that’s either running locally or hosted on a cloud server.
  • mlflow lets you run the MLFlow UI directly from the package, which lightMLFlow does not. Again, lightMLFlow is made to be used with a deployed MLFlow instance, which means that your instance needs to already exist (either on local on on a cloud server).

However, there are also significant advantages to using lightMLFlow over mlflow:

  • lightMLFlow features a friendlier API, with significantly fewer functions, no mlflow::mlflow_* function prefixing (following Tidyverse conventions, lightMLFlow function names are verbs), and improved error handling.
  • lightMLFlow fixes some bugs in mlflow’s API wrapping functions.
  • lightMLFlow is significantly more lightweight than mlflow. It doesn’t depend on httpuv, reticulate, or swagger, and has a more minimal footprint in general.
  • lightMLFlow uses aws.s3 to put and save objects to and from S3, which means you don’t need to have a boto3 install on your machine running your MLFlow code. This is an essential change, as it means that lightMLFlow does not require any Python infrastructure, as opposed to mlflow, which does.
  • mlflow (and, specifically, MLFlow Projects) doesn’t play particularly nicely with renv. The reason for that is that an MLProject file that’s pointed at a Git repo will try to clone and run the code from scratch. But with renv, we like restoring a package cache in CI and baking it into the Docker image that the code lives in so that we don’t need to install all of the R packages the project needs every time we run the project. lightMLFlow hacks its way around this problem by allowing the user to run set_git_tracking_tags(), which tricks the MLFlow REST API into thinking that the code was run from an MLFlow Project even when it wasn’t. This lets you keep your normal (e.g.) renv workflow in place and get the benefit of linked Git commits in the MLFlow UI without actually needing any of the MLProject infrastructure or setup steps.
  • For artifact and model logging, lightMLFlow logs R objects directly so that you don’t need to worry about first saving a file to disk and then copying it to your artifact store.
  • In addition, lightMLFlow allows artifacts to be loaded directly into the R session in one shot, instead of first being saved to disk and then loaded afterwards. This eliminates lines of code and the headache associated with going S3 –> disk –> R by abstracting away the disk reads and writes.
  • In lightMLFlow, create_experiment returns the experiment when one with the specified name already exists, instead of erroring.
  • lightMLFlow adds get_param and get_metric helpers to make it easier to get the most recent value of a metric or param for a run.
  • lightMLFlow leverages some R-ish ways of doing things in reworking log_params and log_metrics, which both take dot args of metrics or params to log, and then call log_batch() on the backend. This results in a far friendlier API than needing to specify a key and value separately and forcing the user to make multiple calls to log_param (e.g.) for each key-value pair. With lightMLFlow, you can do this to log two params: log_params(foo, bar = "baz"), which will log the value of foo as foo (it automagically generates the param name by deparsing the name of the R object), and will log bar as "baz".
  • When logging metrics, lightMLFlow abstracts away timestamps and steps from the user, automatically setting the timestamp to the current time (UTC) and auto-incrementing the step if the metric being logged already exists.
  • If a run errors out, lightMLFlow logs the error as an artifact (a markdown document) to help with the debugging process.

Known Issues / Future Work

  1. Clarify the parameter names for things like path, model_path, etc. since they don’t make much sense right now.
  2. Clarify the difference between save_model and log_model.
  3. Add Azure and GCP artifact stores.

About

A lightweight, opinionated R wrapper for the MLFlow REST API

Resources

Stars

9 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages