Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Pandas 2 assign_in_place() fix with categoricals - #948

Merged
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix
Jul 31, 2025
Merged

Pandas 2 assign_in_place() fix with categoricals#948
jpn-- merged 3 commits into
ActivitySim:mainfrom
i-am-sijia:pd2-categorical-fix

Conversation

@i-am-sijia

Copy link
Copy Markdown
Member

ActivitySim models often call the activitysim.core.util.assign_in_place(df, df2, ...) method to update existing values in df using values from df2, or to add new columns from df2. When performing the former, it calls df.update(df2) to update common columns in place.

In Pandas 2.x, df.update(df2) will raise a TypeError if the common column C in df is of categorical dtype and df2['C'] contains category value(s) that are not present in df['C']'s categories.

This issue was encountered by SANDAG in the trip preprocessor of their airport model, see discussion #946 . None of the existing ActivitySim example models exhibit this issue.

The solution in this PR first makes sure df2['C'] is a categorical type, and then unions the two categoricals before calling df.update(df2).

Comment threadactivitysim/core/util.py Outdated
# when df and df2 column are both categorical, union categories
from pandas.api.types import union_categoricals

uc = union_categoricals([df[c], df2[c]])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I recommend setting sort_categories=True here, which will help make sure that categories populate in a stable order under multiprocessing, which will help with reproducibility/stability, especially when multiprocessing.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

Has this been tested? Can we write a test for it? It looks like it should work, but it also may have unexpected consequences (e.g. from reordering categories). I see the current tests are passing, but they were also passing without this so it's clear our current tests are not exercising the problem being solved.

@jpn--

jpn-- commented Jun 2, 2025

Copy link
Copy Markdown
Member

@i-am-sijia, I just also saw your comment in #946 (comment) . Can we check if bringing "sort_categories=True" to the other existing usage[s] of union_categoricals will fail any of our existing tests?

@i-am-sijiai-am-sijia self-assigned this Jun 27, 2025
@i-am-sijia

Copy link
Copy Markdown
MemberAuthor

@jpn-- I added sort_categories=True to all union_categoricals(). But now I am having second thoughts. Most categorical variables are created with choice alternatives being the pre-defined categories, using the same sort sequence of how the alternatives are defined. Categoricals may not be in the alphabetically order, but they are sorted and include all possible values. This makes categoricals stable under multiprocessing.

cat_type=pd.api.types.CategoricalDtype(
model_spec.columns.tolist() + [""], ordered=False
)
choices=choices.astype(cat_type)

We need union_categoricals() when new category values are added. You are right that under multiprocessing, the sequence of how the new categories are being appended may be unstable. But if we set sort_categories=True during union, it will also change the original sequence of the old category values.

  1. Should we sort all categoricals alphabetically at creation, not just when union is happening? Otherwise you are changing the sequence when union. So that they will be consistent through out.
  2. With sort_categories=True, whenever a new category is append, it re-sorts everything instead of appending, which also impacts the stability. Would this be a problem for Sharrow, i.e., triggers re-compilation?
  3. Side note - In general, I think we do not want to set ordered=True, unless we want to write expressions directly compare categories. I reverted the ordered mode categories. [1] (I don't recall why I specifically made it ordered, I left a note in there in case it causes us trouble in the future)

In terms of unit tests. I was thinking we can start with the following:

  • Concatenating two categorical columns with different categories
  • Overwriting a categorical column with a non-categorical column
  • Overwriting a categorical column with a categorical column with different categories

@jpn--
jpn-- self-requested a review July 17, 2025 18:14
@jpn--

Copy link
Copy Markdown
Member

Maybe we are getting too far ahead of ourselves.

I agree that most categoricals are fully enumerated at creation time, in some logical order that may be meaningful to modelers (e.g. driving, then transit, then nonmotorized). And, most categoricals in ActivitySim are logically unordered.

For sharrow, any change in a categorical data type (adding categories, removing unused categories, reordering them, making them "ordered" when they were previously just nominal), literally any change at all will result in recompiling. As long as we get stable sort ordering most of the time, we can survive with corner cases that don't have stable order (e.g. the sample is too small and some unusual tour type doesn't always get included).

So, maybe we just merge this to fix the bug encountered as reported at the top, and leave a more complete solution to efficiently handling categoricals until ActivitySim 2.x?

@jpn--
jpn-- merged commit 58dc347 into ActivitySim:mainJul 31, 2025
17 checks passed
@i-am-sijia
i-am-sijia deleted the pd2-categorical-fix branch October 17, 2025 20:50
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@i-am-sijia@jpn--