Skip to content

[SPARK-57991][CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26 - #57057

Closed
HyukjinKwon wants to merge 1 commit into
apache:masterfrom
HyukjinKwon:deflake-barrier-macos26
Closed

[SPARK-57991][CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26#57057
HyukjinKwon wants to merge 1 commit into
apache:masterfrom
HyukjinKwon:deflake-barrier-macos26

Conversation

@HyukjinKwon

@HyukjinKwonHyukjinKwon commented Jul 6, 2026

Copy link
Copy Markdown
Member

What changes were proposed in this pull request?

Widen the post-sync wall-clock skew tolerance in BarrierTaskContextSuite from a hard-coded
<= 1000 (ms) to a documented maxSyncSkewMs = 3000 constant, applied to the three tests that
assert how closely tasks finish a barrier() / allGather() global sync:

  • global sync by barrier() call
  • successively sync with allGather and barrier
  • support multiple barrier() call within a single task

Why are the changes needed?

The Build / Maven (Scala 2.13, JDK 21, MacOS-26) scheduled workflow has been failing on every
recent run. One of the reproducing failures is:

BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)

These tests capture System.currentTimeMillis() in each task immediately after a
barrier/allGather returns and assert times.max - times.min <= 1000. That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.

BarrierTaskContext.runBarrier waits for the coordinator RPC in a Thread.sleep(1000) loop
(BarrierTaskContext.scala:102),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.

Evidence gathered:

The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.

Scope note

This PR fixes only the BarrierTaskContextSuite flake. The same workflow also exhibits other,
independent flakes (Kafka stress test for failOnDataLoss=false topic-deletion timeout, mllib GMM
FP-tolerance on arm64, StateStoreSuite.maintenance timeout). Those are out of scope here and
would be separate changes. (The earlier AF_UNIX path too long UDF-worker flake was already fixed
upstream by SPARK-57949.)

Does this PR introduce any user-facing change?

No. Test-only.

How was this patch tested?

Locally on macOS 26 arm64 / JDK 21:
build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite
— all 13 tests pass; scalastyle clean.

Also validated end-to-end on the Build / Maven (Scala 2.13, JDK 21, MacOS-26) workflow,
running the same core matrix job on the identical macos-26 arm64 runner image:

core,launcher,… job (runs BarrierTaskContextSuite)Result
Before (apache/spark master, without this fix)runs/28790673283 · job 853790295521078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)Tests: succeeded 4103, failed 1
After (this branch)runs/28829214565 · job 85505367410Tests: succeeded 4104, failed 0; the three barrier tests all pass

(To run the guarded workflow on the fork, its if: github.repository == 'apache/spark'
condition was relaxed on a throwaway commit that is not part of this PR — this PR is the
single-file test change only.)

Was this patch authored or co-authored using generative AI tooling?

Yes, drafted with assistance from Claude.

@HyukjinKwon
HyukjinKwonforce-pushed the deflake-barrier-macos26 branch from 8cef6cc to 35a5a4cCompareJuly 7, 2026 04:14
@HyukjinKwon
HyukjinKwon marked this pull request as ready for review July 7, 2026 04:15
@HyukjinKwon
HyukjinKwonforce-pushed the deflake-barrier-macos26 branch from 35a5a4c to 744da65CompareJuly 7, 2026 04:27
@HyukjinKwonHyukjinKwon changed the title [DO-NOT-MERGE][CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26[CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26Jul 7, 2026
@dongjoon-hyun

Copy link
Copy Markdown
Member

Could you file a JIRA issue, @HyukjinKwon ?

… skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite`
from a hard-coded `<= 1000` (ms) to a documented `maxSyncSkewMs = 3000`
constant, applied to the three tests that assert how closely tasks finish a
`barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
These tests capture `System.currentTimeMillis()` in each task right after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. That
skew is bounded below by the barrier client's own polling granularity, not by
sync correctness: `BarrierTaskContext.runBarrier` waits for the coordinator
RPC in a `Thread.sleep(1000)` loop, so two tasks can observe the release up to
~1s apart before any thread-scheduling or GC jitter is added. On a busy CI host
this overflows the 1000ms bound with no real bug.
Evidence:
- macOS-26 arm64 scheduled run failed with `1078 was not less than or equal to
1000 (BarrierTaskContextSuite.scala:122)`.
- SPARK-49983 (2024) previously observed `1038` and only halved the pre-barrier
sleep, which reduces arrival spread but cannot reduce the poll-granularity
skew, so the flake persists on slower runners.
- Reproduced locally on macOS 26 arm64 / JDK 21: measured post-sync skew up to
994ms while completely idle, i.e. essentially zero margin under the old bound.
The new bound (twice the 1s poll interval plus a jitter allowance) keeps the
test's intent — tasks stay loosely in lockstep and none races an entire extra
sync ahead — without coupling to a razor-thin wall-clock margin.
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
on macOS 26 arm64 with JDK 21: all 13 tests pass, scalastyle clean.
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
@HyukjinKwon
HyukjinKwonforce-pushed the deflake-barrier-macos26 branch from 744da65 to c2f1bdeCompareJuly 7, 2026 06:01
@HyukjinKwonHyukjinKwon changed the title [CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26[SPARK-57991][CORE][TESTS] Deflake BarrierTaskContextSuite wall-clock skew assertions on macOS-26Jul 7, 2026
HyukjinKwon added a commit that referenced this pull request Jul 7, 2026
…skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite` from a hard-coded
`<= 1000` (ms) to a documented `maxSyncSkewMs = 3000` constant, applied to the three tests that
assert how closely tasks finish a `barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
The `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` scheduled workflow has been failing on every
recent run. One of the reproducing failures is:
```
BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)
```
These tests capture `System.currentTimeMillis()` in each task immediately after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. **That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.**
`BarrierTaskContext.runBarrier` waits for the coordinator RPC in a `Thread.sleep(1000)` loop
([`BarrierTaskContext.scala:102`](https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/BarrierTaskContext.scala#L102)),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.
Evidence gathered:
- SPARK-49983 (2024, PR #48487) previously saw `1038` and only *halved the pre-barrier sleep*.
That reduces arrival spread but cannot reduce poll-granularity skew, so the flake persists on
slower runners like `macos-26` arm64.
- Reproduced locally on **macOS 26 arm64 / JDK 21** (the same OS/arch as the failing runner):
measured post-sync skew up to **994 ms while the machine was idle** — i.e. essentially zero
margin under the old `1000` bound, so any CI load tips it over.
The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.
### Scope note
This PR fixes only the `BarrierTaskContextSuite` flake. The same workflow also exhibits other,
independent flakes (Kafka `stress test for failOnDataLoss=false` topic-deletion timeout, mllib GMM
FP-tolerance on arm64, `StateStoreSuite.maintenance` timeout). Those are out of scope here and
would be separate changes. (The earlier `AF_UNIX path too long` UDF-worker flake was already fixed
upstream by SPARK-57949.)
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
Locally on macOS 26 arm64 / JDK 21:
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
— all 13 tests pass; scalastyle clean.
Also validated end-to-end on the `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` workflow,
running the same `core` matrix job on the identical `macos-26` arm64 runner image:
| | `core,launcher,…` job (runs `BarrierTaskContextSuite`) | Result |
|---|---|---|
| **Before** (`apache/spark` master, without this fix) | [runs/28790673283 · job 85379029552](https://github.com/apache/spark/actions/runs/28790673283/job/85379029552) | ❌ `1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)` — `Tests: succeeded 4103, failed 1` |
| **After** (this branch) | [runs/28829214565 · job 85505367410](https://github.com/HyukjinKwon/spark/actions/runs/28829214565/job/85505367410) | ✅ `Tests: succeeded 4104, failed 0`; the three barrier tests all pass |
(To run the guarded workflow on the fork, its `if: github.repository == 'apache/spark'`
condition was relaxed on a throwaway commit that is **not** part of this PR — this PR is the
single-file test change only.)
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
Closes#57057 from HyukjinKwon/deflake-barrier-macos26.
Authored-by: Hyukjin Kwon <gurwls223@apache.org>
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
(cherry picked from commit e887fd5)
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
HyukjinKwon added a commit that referenced this pull request Jul 7, 2026
…skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite` from a hard-coded
`<= 1000` (ms) to a documented `maxSyncSkewMs = 3000` constant, applied to the three tests that
assert how closely tasks finish a `barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
The `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` scheduled workflow has been failing on every
recent run. One of the reproducing failures is:
```
BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)
```
These tests capture `System.currentTimeMillis()` in each task immediately after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. **That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.**
`BarrierTaskContext.runBarrier` waits for the coordinator RPC in a `Thread.sleep(1000)` loop
([`BarrierTaskContext.scala:102`](https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/BarrierTaskContext.scala#L102)),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.
Evidence gathered:
- SPARK-49983 (2024, PR #48487) previously saw `1038` and only *halved the pre-barrier sleep*.
That reduces arrival spread but cannot reduce poll-granularity skew, so the flake persists on
slower runners like `macos-26` arm64.
- Reproduced locally on **macOS 26 arm64 / JDK 21** (the same OS/arch as the failing runner):
measured post-sync skew up to **994 ms while the machine was idle** — i.e. essentially zero
margin under the old `1000` bound, so any CI load tips it over.
The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.
### Scope note
This PR fixes only the `BarrierTaskContextSuite` flake. The same workflow also exhibits other,
independent flakes (Kafka `stress test for failOnDataLoss=false` topic-deletion timeout, mllib GMM
FP-tolerance on arm64, `StateStoreSuite.maintenance` timeout). Those are out of scope here and
would be separate changes. (The earlier `AF_UNIX path too long` UDF-worker flake was already fixed
upstream by SPARK-57949.)
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
Locally on macOS 26 arm64 / JDK 21:
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
— all 13 tests pass; scalastyle clean.
Also validated end-to-end on the `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` workflow,
running the same `core` matrix job on the identical `macos-26` arm64 runner image:
| | `core,launcher,…` job (runs `BarrierTaskContextSuite`) | Result |
|---|---|---|
| **Before** (`apache/spark` master, without this fix) | [runs/28790673283 · job 85379029552](https://github.com/apache/spark/actions/runs/28790673283/job/85379029552) | ❌ `1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)` — `Tests: succeeded 4103, failed 1` |
| **After** (this branch) | [runs/28829214565 · job 85505367410](https://github.com/HyukjinKwon/spark/actions/runs/28829214565/job/85505367410) | ✅ `Tests: succeeded 4104, failed 0`; the three barrier tests all pass |
(To run the guarded workflow on the fork, its `if: github.repository == 'apache/spark'`
condition was relaxed on a throwaway commit that is **not** part of this PR — this PR is the
single-file test change only.)
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
Closes#57057 from HyukjinKwon/deflake-barrier-macos26.
Authored-by: Hyukjin Kwon <gurwls223@apache.org>
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
(cherry picked from commit e887fd5)
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
HyukjinKwon added a commit that referenced this pull request Jul 7, 2026
…skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite` from a hard-coded
`<= 1000` (ms) to a documented `maxSyncSkewMs = 3000` constant, applied to the three tests that
assert how closely tasks finish a `barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
The `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` scheduled workflow has been failing on every
recent run. One of the reproducing failures is:
```
BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)
```
These tests capture `System.currentTimeMillis()` in each task immediately after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. **That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.**
`BarrierTaskContext.runBarrier` waits for the coordinator RPC in a `Thread.sleep(1000)` loop
([`BarrierTaskContext.scala:102`](https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/BarrierTaskContext.scala#L102)),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.
Evidence gathered:
- SPARK-49983 (2024, PR #48487) previously saw `1038` and only *halved the pre-barrier sleep*.
That reduces arrival spread but cannot reduce poll-granularity skew, so the flake persists on
slower runners like `macos-26` arm64.
- Reproduced locally on **macOS 26 arm64 / JDK 21** (the same OS/arch as the failing runner):
measured post-sync skew up to **994 ms while the machine was idle** — i.e. essentially zero
margin under the old `1000` bound, so any CI load tips it over.
The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.
### Scope note
This PR fixes only the `BarrierTaskContextSuite` flake. The same workflow also exhibits other,
independent flakes (Kafka `stress test for failOnDataLoss=false` topic-deletion timeout, mllib GMM
FP-tolerance on arm64, `StateStoreSuite.maintenance` timeout). Those are out of scope here and
would be separate changes. (The earlier `AF_UNIX path too long` UDF-worker flake was already fixed
upstream by SPARK-57949.)
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
Locally on macOS 26 arm64 / JDK 21:
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
— all 13 tests pass; scalastyle clean.
Also validated end-to-end on the `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` workflow,
running the same `core` matrix job on the identical `macos-26` arm64 runner image:
| | `core,launcher,…` job (runs `BarrierTaskContextSuite`) | Result |
|---|---|---|
| **Before** (`apache/spark` master, without this fix) | [runs/28790673283 · job 85379029552](https://github.com/apache/spark/actions/runs/28790673283/job/85379029552) | ❌ `1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)` — `Tests: succeeded 4103, failed 1` |
| **After** (this branch) | [runs/28829214565 · job 85505367410](https://github.com/HyukjinKwon/spark/actions/runs/28829214565/job/85505367410) | ✅ `Tests: succeeded 4104, failed 0`; the three barrier tests all pass |
(To run the guarded workflow on the fork, its `if: github.repository == 'apache/spark'`
condition was relaxed on a throwaway commit that is **not** part of this PR — this PR is the
single-file test change only.)
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
Closes#57057 from HyukjinKwon/deflake-barrier-macos26.
Authored-by: Hyukjin Kwon <gurwls223@apache.org>
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
(cherry picked from commit e887fd5)
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
HyukjinKwon added a commit that referenced this pull request Jul 7, 2026
…skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite` from a hard-coded
`<= 1000` (ms) to a documented `maxSyncSkewMs = 3000` constant, applied to the three tests that
assert how closely tasks finish a `barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
The `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` scheduled workflow has been failing on every
recent run. One of the reproducing failures is:
```
BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)
```
These tests capture `System.currentTimeMillis()` in each task immediately after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. **That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.**
`BarrierTaskContext.runBarrier` waits for the coordinator RPC in a `Thread.sleep(1000)` loop
([`BarrierTaskContext.scala:102`](https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/BarrierTaskContext.scala#L102)),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.
Evidence gathered:
- SPARK-49983 (2024, PR #48487) previously saw `1038` and only *halved the pre-barrier sleep*.
That reduces arrival spread but cannot reduce poll-granularity skew, so the flake persists on
slower runners like `macos-26` arm64.
- Reproduced locally on **macOS 26 arm64 / JDK 21** (the same OS/arch as the failing runner):
measured post-sync skew up to **994 ms while the machine was idle** — i.e. essentially zero
margin under the old `1000` bound, so any CI load tips it over.
The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.
### Scope note
This PR fixes only the `BarrierTaskContextSuite` flake. The same workflow also exhibits other,
independent flakes (Kafka `stress test for failOnDataLoss=false` topic-deletion timeout, mllib GMM
FP-tolerance on arm64, `StateStoreSuite.maintenance` timeout). Those are out of scope here and
would be separate changes. (The earlier `AF_UNIX path too long` UDF-worker flake was already fixed
upstream by SPARK-57949.)
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
Locally on macOS 26 arm64 / JDK 21:
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
— all 13 tests pass; scalastyle clean.
Also validated end-to-end on the `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` workflow,
running the same `core` matrix job on the identical `macos-26` arm64 runner image:
| | `core,launcher,…` job (runs `BarrierTaskContextSuite`) | Result |
|---|---|---|
| **Before** (`apache/spark` master, without this fix) | [runs/28790673283 · job 85379029552](https://github.com/apache/spark/actions/runs/28790673283/job/85379029552) | ❌ `1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)` — `Tests: succeeded 4103, failed 1` |
| **After** (this branch) | [runs/28829214565 · job 85505367410](https://github.com/HyukjinKwon/spark/actions/runs/28829214565/job/85505367410) | ✅ `Tests: succeeded 4104, failed 0`; the three barrier tests all pass |
(To run the guarded workflow on the fork, its `if: github.repository == 'apache/spark'`
condition was relaxed on a throwaway commit that is **not** part of this PR — this PR is the
single-file test change only.)
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
Closes#57057 from HyukjinKwon/deflake-barrier-macos26.
Authored-by: Hyukjin Kwon <gurwls223@apache.org>
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
(cherry picked from commit e887fd5)
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
HyukjinKwon added a commit that referenced this pull request Jul 7, 2026
…skew assertions on macOS-26
### What changes were proposed in this pull request?
Widen the post-sync wall-clock skew tolerance in `BarrierTaskContextSuite` from a hard-coded
`<= 1000` (ms) to a documented `maxSyncSkewMs = 3000` constant, applied to the three tests that
assert how closely tasks finish a `barrier()` / `allGather()` global sync:
- `global sync by barrier() call`
- `successively sync with allGather and barrier`
- `support multiple barrier() call within a single task`
### Why are the changes needed?
The `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` scheduled workflow has been failing on every
recent run. One of the reproducing failures is:
```
BarrierTaskContextSuite:
- successively sync with allGather and barrier *** FAILED ***
1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)
```
These tests capture `System.currentTimeMillis()` in each task immediately after a
barrier/allGather returns and assert `times.max - times.min <= 1000`. **That skew is bounded below
by the barrier client's own polling granularity, not by sync correctness.**
`BarrierTaskContext.runBarrier` waits for the coordinator RPC in a `Thread.sleep(1000)` loop
([`BarrierTaskContext.scala:102`](https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/BarrierTaskContext.scala#L102)),
so two tasks can observe the release up to ~1s apart from the poll interval alone — before any
thread-scheduling or GC jitter on a busy CI host is added.
Evidence gathered:
- SPARK-49983 (2024, PR #48487) previously saw `1038` and only *halved the pre-barrier sleep*.
That reduces arrival spread but cannot reduce poll-granularity skew, so the flake persists on
slower runners like `macos-26` arm64.
- Reproduced locally on **macOS 26 arm64 / JDK 21** (the same OS/arch as the failing runner):
measured post-sync skew up to **994 ms while the machine was idle** — i.e. essentially zero
margin under the old `1000` bound, so any CI load tips it over.
The new bound (twice the 1 s poll interval plus a jitter allowance) preserves the test's intent —
tasks stay loosely in lockstep and none can race an entire extra sync ahead — without coupling the
assertion to a razor-thin wall-clock margin.
### Scope note
This PR fixes only the `BarrierTaskContextSuite` flake. The same workflow also exhibits other,
independent flakes (Kafka `stress test for failOnDataLoss=false` topic-deletion timeout, mllib GMM
FP-tolerance on arm64, `StateStoreSuite.maintenance` timeout). Those are out of scope here and
would be separate changes. (The earlier `AF_UNIX path too long` UDF-worker flake was already fixed
upstream by SPARK-57949.)
### Does this PR introduce _any_ user-facing change?
No. Test-only.
### How was this patch tested?
Locally on macOS 26 arm64 / JDK 21:
`build/mvn -pl core test -DwildcardSuites=org.apache.spark.scheduler.BarrierTaskContextSuite`
— all 13 tests pass; scalastyle clean.
Also validated end-to-end on the `Build / Maven (Scala 2.13, JDK 21, MacOS-26)` workflow,
running the same `core` matrix job on the identical `macos-26` arm64 runner image:
| | `core,launcher,…` job (runs `BarrierTaskContextSuite`) | Result |
|---|---|---|
| **Before** (`apache/spark` master, without this fix) | [runs/28790673283 · job 85379029552](https://github.com/apache/spark/actions/runs/28790673283/job/85379029552) | ❌ `1078 was not less than or equal to 1000 (BarrierTaskContextSuite.scala:122)` — `Tests: succeeded 4103, failed 1` |
| **After** (this branch) | [runs/28829214565 · job 85505367410](https://github.com/HyukjinKwon/spark/actions/runs/28829214565/job/85505367410) | ✅ `Tests: succeeded 4104, failed 0`; the three barrier tests all pass |
(To run the guarded workflow on the fork, its `if: github.repository == 'apache/spark'`
condition was relaxed on a throwaway commit that is **not** part of this PR — this PR is the
single-file test change only.)
### Was this patch authored or co-authored using generative AI tooling?
Yes, drafted with assistance from Claude.
Closes#57057 from HyukjinKwon/deflake-barrier-macos26.
Authored-by: Hyukjin Kwon <gurwls223@apache.org>
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
(cherry picked from commit e887fd5)
Signed-off-by: Hyukjin Kwon <hyukjin.kwon@databricks.com>
@HyukjinKwon

Copy link
Copy Markdown
MemberAuthor

Merge Summary:

Posted by merge_spark_pr.py

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@HyukjinKwon@dongjoon-hyun