Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Update to use pandas v2.* - #932

Merged
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2
May 20, 2025
Merged

Update to use pandas v2.*#932
jpn-- merged 34 commits into
ActivitySim:mainfrom
driftlesslabs:pandas-2

Conversation

@jpn--

@jpn--jpn-- commented Mar 18, 2025

Copy link
Copy Markdown
Member

Addresses #794.

The update from pandas 1.x to 2.x introduces a number of small but material changes that affect ActivitySim:

  • DataFrame Index objects are all one class with different datatypes, instead of being different classes (e.g. there is no more Int64Index class).
  • The read_csv function by default now interprets "None" as a missing value (i.e. NaN) instead of being the Python object None.
  • The groupby operation, when applied to categorical data, now sorts the categories in the result unless told not to (resulting in different order of rows in outputs for some operations).
  • A simple df.join() also potentially sorts the resulting rows differently unless an explicit sort argument is given.
  • Index objects no longer can be checked as is_monotonic but instead need is_monotonic_increasing.
  • The handling of dtypes appears to have improved in some instances, where dtypes used to be promoted by some operations now they are not (e.g. variables that are originally int16 used to become int64 after some operations and now they don't).

This pull request includes several changes across multiple files to address these pandas changes. The most important changes include modifications to sorting operations, error handling in logging, and the introduction of a new fast_eval function to optimize DataFrame evaluations, because the regular pandas.eval has some significant performance degradations.

Data Handling Improvements:

Error Handling Enhancements:

Evaluation Process Optimization:

  • Introduced fast_eval function in activitysim/core/fast_eval.py to optimize DataFrame evaluations by handling special characters in column names and improving performance.
  • Updated references to df.eval in activitysim/core/interaction_simulate.py and activitysim/core/simulate.py to use fast_eval for better performance and consistency. [1][2][3][4][5]

Miscellaneous Changes:

jpn-- added 29 commits March 18, 2024 18:11
# Conflicts:
#	conda-environments/activitysim-dev.yml
#	conda-environments/github-actions-tests.yml
# Conflicts:
#	.github/workflows/core_tests.yml
#	activitysim/abm/models/trip_departure_choice.py
#	activitysim/abm/models/vehicle_allocation.py
#	activitysim/examples/prototype_mtc_extended/test/prototype_mtc_extended_reference_pipeline.zip
#	conda-environments/activitysim-dev.yml
#	conda-environments/docbuild.yml
#	conda-environments/github-actions-tests.yml
#	pyproject.toml
# Conflicts:
#	conda-environments/docbuild.yml
@jpn--jpn-- changed the title Pandas 2Update to use pandas v2.*Mar 19, 2025
@jpn--

Copy link
Copy Markdown
MemberAuthor

The changes I have made in this new branch have greatly improved runtime performance while using pandas 2.x.

non-sharrow test timings for pandas 1.x:

58.60s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
53.71s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
53.66s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
53.23s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

first attempt non-sharrow test timings for pandas 2.x (#838):

148.50s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
148.14s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
147.83s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode
140.09s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp

revised non-sharrow test timings for pandas 2.x (this PR, #932):

65.06s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_mp
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_chunkless
58.10s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc
57.38s call activitysim/examples/prototype_mtc/test/test_mtc.py::test_mtc_recode

We can see that there is still a modest runtime cost to using pandas 2.x, on the order of 10% slower, but nowhere near the cost of the first attempt, which was ~200% slower. Achieving no runtime penalty appears to be possible, but it would require accessing non-public pandas functions which might break in the future, see here

Note all of these runtime issues are exclusively non-sharrow, as sharrow evaluation completely bypasses the pandas.eval function that is the source of our problem.

@jpn--
jpn-- requested a review from i-am-sijiaMarch 19, 2025 14:30
@jpn--jpn-- mentioned this pull request Mar 19, 2025
@jpn--
jpn-- requested a review from CopilotMarch 31, 2025 19:51

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR updates the codebase for compatibility with pandas v2, addressing changes in DataFrame indexing, evaluation methods, and error handling while improving performance with a new fast_eval function.

  • Replaces several instances of DataFrame.eval with a custom fast_eval function to enhance performance and handle special characters in column names.
  • Introduces sorting and index reset adjustments across multiple functions to ensure data consistency.
  • Updates resource handling to use importlib.resources, and improves error logging in various modules.

Reviewed Changes

Copilot reviewed 26 out of 27 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
activitysim/examples/placeholder_sandag/test/test_sandag.pyAdded new test for local compute with updated configs.
activitysim/estimation/larch/simple_simulate.pyReplaced DataFrame.eval with fast_eval for evaluation.
activitysim/estimation/larch/scheduling.pySwitched to fast_eval to optimize evaluation.
activitysim/core/workflow/state.pyEnhanced error handling when creating pa.Table from DataFrames.
activitysim/core/util.pyAdjusted index type checking to align with pandas v2.
activitysim/core/test/_tools.pyImproved error reporting with exception details.
activitysim/core/simulate.pyUpdated DataFrame evaluation to use fast_eval.
activitysim/core/los.pyModified type conversions to prevent numeric overflow.
activitysim/core/interaction_simulate.pyReplaced df.eval with fast_eval for consistency and performance.
activitysim/core/fast_eval.pyIntroduced fast_eval function to optimize DataFrame evaluations.
activitysim/core/assign.pyUpdated CSV reading with explicit na_values for pandas v2 behavior.
activitysim/cli/create.pyModernized resource handling using importlib.resources.
activitysim/abm/models/vehicle_allocation.pyEnforced correct dtype conversion for vehicle choices.
activitysim/abm/models/util/school_escort_tours_trips.pyAdded reset_index(drop=True) to ensure consistent indexing.
activitysim/abm/models/trip_departure_choice.pyUpdated monotonic index check to is_monotonic_increasing.
activitysim/abm/models/school_escorting.pyReset index on escort_bundles to maintain data integrity.
activitysim/abm/models/input_checker.pyEnhanced error logging with exception details in validators.
activitysim/abm/models/disaggregate_accessibility.pyAdded sorting after joins to ensure template consistency.
.github/workflows/core_tests.ymlUpdated CI branch references to reflect pandas v2 changes.
Files not reviewed (1)
  • activitysim/examples/prototype_mtc_extended/configs/trip_mode_choice_annotate_trips_preprocessor.csv: Language not supported
Comments suppressed due to low confidence (1)

activitysim/cli/create.py:183

  • Using a context-managed path from importlib.resources.as_file may lead to unexpected behavior when used with glob.glob. Ensure that the returned path is valid for directory globbing in all environments.
for asset_path in glob.glob(str(pth)):

Comment threadactivitysim/core/fast_eval.py
@JoeJimFlood

Copy link
Copy Markdown
Contributor

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

@i-am-sijia

Copy link
Copy Markdown
Member

Have any runtime comparisons been done with the full sandag-abm3-example or anything larger than the 25-zone prototype_mtc example? I'm concerned about nonlinearity in the relationship between the size of a model and the runtime.

Hi @JoeJimFlood, that is a valid concern. As I am reviewing this PR, I can perform the run time test with the full size example SANDAG.

@i-am-sijia

Copy link
Copy Markdown
Member

I ran the full size sandag-abm3-example with this PR (e.g., pandas 2.x) and the main branch (e.g., pandas 1.4) and would like to share some quick initial reports. I do have some other comments which I will post separately.

For both runs, I used:

  • Sharrow: False
  • multiprocess: True
  • num_processes: 5
  • explicit_chunk: 0.2 (for select components)

The run time is almost the same for the two runs, see below. Pandas 2.x runs faster for some components but not the others, e.g., it's faster in non-mandatory tour scheduling but slower in mandatory tour scheduling, which could be just runtime noise. In total, pandas 2.x took ~5 mins longer which is probably negligible. The total run time is comparable to the run time I reported during Phase 9A: ActivitySim/sandag-abm3-example#9 (comment). This PR does shorten the run time for pandas 2.0 as it promised.

I have not checked if the results of the two run are the same, I will check that.

image

@jpn--

Copy link
Copy Markdown
MemberAuthor

I see long sequences where one version is like ~10% faster, or ~10% slower, in sequential contiguous blocks across fairly disparate component types. This strongly suggests much of the runtime differences are external noise from other subprocesses or other issues (e.g. the server got too hot and throttled the compute for a couple minutes).

@JoeJimFlood

Copy link
Copy Markdown
Contributor

Thanks for running that @i-am-sijia! The 1.5% increase in using Pandas v2 vs v1 is encouraging to see.

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm curious about the implications on dependence lock and expression rules. With this PR, ActivitySim will use its own fast_eval() and some rewrite version of internal pandas methods until pandas releases an official version (say pandas 3.0) that fixes our problem at hand. Is the plan for us to be locked with pandas 2.2 and fast_eval.py until then? Otherwise we are adding an overhead to maintain the compatibility of fast_eval() when we'd use pandas >2.2. In terms of expression rules, I saw the comment related to pd.Series in fast_eval.py, was wondering if we should proactively alert users about that.

Comment threadactivitysim/core/interaction_simulate.py
Comment threadconda-environments/activitysim-dev-base.yml
Comment thread.github/workflows/core_tests.yml
Comment threadactivitysim/core/fast_eval.py

@i-am-sijiai-am-sijia left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for responding to my comments. I'll approve this PR.

@jpn--
jpn-- merged commit 146c7ff into ActivitySim:mainMay 20, 2025
@jpn--jpn-- mentioned this pull request Aug 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@jpn--@JoeJimFlood@i-am-sijia