Skip to content

Turn on Helix Job Monitor - #38751

Draft
wtgodbe wants to merge 2 commits into
mainfrom
feature/helix-job-monitor
Draft

Turn on Helix Job Monitor#38751
wtgodbe wants to merge 2 commits into
mainfrom
feature/helix-job-monitor

Conversation

@wtgodbe

Copy link
Copy Markdown
Member

Enable Arcade's Helix Job Monitor for the public and internal test pipelines. Helix submission jobs can now release their agents after queueing work, while the monitor publishes test results and owns the final Helix status.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@wtgodbe
wtgodbe marked this pull request as ready for review August 5, 2026 17:25
CopilotAI review requested due to automatic review settings August 5, 2026 17:25
@wtgodbe

Copy link
Copy Markdown
MemberAuthor

@AndriySvyryd PTAL

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Enables Arcade’s Helix Job Monitor in EF Core’s public and internal Azure DevOps pipelines so Helix submission jobs can stop after queueing, while a dedicated monitor job publishes test results and drives the final Helix status.

Changes:

  • Add Microsoft.DotNet.Helix.JobMonitor dependency/version plumbing and pin the tool via .config/dotnet-tools.json.
  • Enable Helix Job Monitor behavior in eng/helix.proj when SYSTEM_ACCESSTOKEN is available.
  • Add the helix-job-monitor.yml job template to both public and internal pipelines.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.

Show a summary per file
FileDescription
eng/Version.Details.xmlAdds Helix Job Monitor dependency tracking entry.
eng/Version.Details.propsIntroduces Helix Job Monitor version properties alongside other dotnet-dotnet dependencies.
eng/helix.projTurns on Helix Job Monitor mode when SYSTEM_ACCESSTOKEN is set.
azure-pipelines-public.ymlAdds Helix Job Monitor job template to the public pipeline (currently with an indentation issue).
azure-pipelines-internal-tests.ymlAdds Helix Job Monitor job template to the internal test pipeline and passes helixAccessToken.
.config/dotnet-tools.jsonPins the dotnet-helix-job-monitor tool version for dotnet tool restore.

env:
HelixAccessToken: $(_HelixAccessToken)
SYSTEM_ACCESSTOKEN: $(System.AccessToken)
- template: /eng/common/core-templates/job/helix-job-monitor.yml
- template: /eng/common/core-templates/job/helix-job-monitor.yml@self
parameters:
helixAccessToken: $(HelixApiAccessToken)
- stage: validate

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You need to change the validate logic to also check whether the corresponding Helix monitor job succeeded

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally, we should check the result of each individual Helix job to preserve the validation logic, though I am not sure whether this is currently possible to do here (feature request?)
Otherwise, add $helixJobMonitorResult to each item in $groupResults instead of failing outright

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally, we should check the result of each individual Helix job to preserve the validation logic

What would be the benefit of this? The helix jobs no longer depend on the test results, they just send the tests off and then report green. If one fails, the monitor will fail too

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What would be the benefit of this? The helix jobs no longer depend on the test results, they just send the tests off and then report green. If one fails, the monitor will fail too

Right, the benefit would be from checking the helix monitoring jobs for test failures specific to that leg

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the benefit would be from checking the helix monitoring jobs for test failures specific to that leg

But the individual legs don't fail when there are test failures - as soon as the tests are sent to helix, they complete w/ success. Only the helix monitor job will ever fail for test failures.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only the helix monitor job will ever fail for test failures.

Yes and we need to make that failure more granular, so that we can continue to check only the relevant failures in the validation groups.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings August 5, 2026 17:34

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 6 out of 6 changed files in this pull request and generated no new comments.

@wtgodbe
wtgodbe marked this pull request as draft August 6, 2026 17:36
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@wtgodbe@AndriySvyryd