Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add options to handle larger dataset for location models - #687

Merged
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch
Feb 8, 2024
Merged

Add options to handle larger dataset for location models#687
jpn-- merged 6 commits into
ActivitySim:developfrom
TransLinkForecasting:location_estimation_patch

Conversation

@bwentl

Copy link
Copy Markdown
Contributor

Currently, the location_choice.py used by from activitysim.estimation.larch import component_model cannot load the estimation data in reasonable length of time. See issue #686.

Description of the problem

The main reason why that is the case is that activitysim uses a chooser-variable table by alternatives – “cv” (zone 1, 2, etc. as columns, and attributes as rows), whereas larch uses a idca table by attributes – “ca” (dist, accessibility, etc. attributes as columns and zones as rows). In order to work with larch, activitysim converts the table format using the cv_to_ca function. However, this function does not scale well with large tables, and we found some modification is required to process the data.

Changes proposed

Here is a summary of the changes we did to location_choice.py to ensure the workplace location model data can be processed faster:

  • Added the conversion of csv data to feather to speed up subsequent reads of the estimation data bundle – note that we did ran into issue with loading versions of csv data that have mixed type in a column when using feather, in that case, set alt_values_to_feather=False (which is the default) prevent errors. This however should not happen if location model specifications are properly written.
  • Add chunking to the processing of the alternatives data, we typically use chunking size of 50000 rows of the cv-format table, but you might need to adjust this depending the spec of your machine or the requirement of your model. The default of this is chunking_size=None, which fall back to the original method of processing the entire dataset.

Usage

In the estimation notebook for school or workplace location models, modify the way model and data are loaded with activitysim.estimation.larch.component_model:

modelname = "workplace_location"
from activitysim.estimation.larch import component_model
# set chunking size for x_ca. Note that this parameter has no effect if "{name}_x_ca.pkl" cache is present.
chunking_size = 50000
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)

bwentl added 3 commits June 23, 2023 09:51
use feather and pickle for faster loading
split x_ca processing into idca using chunking_size
and update comments
default behavior does not change, to enable use of feather
and the use of chunking, do this:
# load from original
model, data = component_model(modelname,
return_data=True,
alt_values_to_feather=True,
chunking_size=chunking_size)
@bwentl
bwentlforce-pushed the location_estimation_patch branch from b8beb66 to 425467eCompareNovember 2, 2023 00:01
@jpn--
jpn-- changed the base branch from main to developJanuary 31, 2024 18:53
@bwentl

Copy link
Copy Markdown
ContributorAuthor

@jpn-- I have brought my branch up to date with develop now, and the merge should work now.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@bwentl

bwentl commented Feb 7, 2024

Copy link
Copy Markdown
ContributorAuthor

@bwentl Does this work for all location/destination models? Non-mandatory tour destination, for example? Thanks!

@stefancoe I believe so. The model type and the list of output data are the same for mandatory and non-mandatory tour destinations, so it should work. The only thing that needs to be changed is the modelname when you load the data for estimation.

@stefancoe

Copy link
Copy Markdown
Contributor

@bwentl Thanks- I'll try it out.

@jpn--
jpn-- merged commit 6a20954 into ActivitySim:developFeb 8, 2024
@jpn--jpn-- mentioned this pull request Feb 14, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@bwentl@stefancoe@jpn--