Uh oh!
There was an error while loading. Please reload this page.
[SPARK-28869][CORE] Roll over event log files - #25670
Conversation
HeartSaVioR
commented
Sep 3, 2019
Fixing it would be simple - clone the properties and include into |
HeartSaVioR
commented
Sep 3, 2019
cc. @felixcheung (as shepherd of SPARK-28594) @vanzin@squito@gengliangwang@dongjoon-hyun also cc. to @Ngone51 as might be interested on this. |
SparkQA
commented
Sep 3, 2019
Test build #110065 has finished for PR 25670 at commit
|
SparkQA
commented
Sep 4, 2019
Test build #110068 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 4, 2019
This is also occurring with current master branch. You can reproduce it with below query in spark-shell, with continuously pushing records to topic1 and topic2. |
SparkQA
commented
Sep 4, 2019
Test build #110073 has finished for PR 25670 at commit
|
felixcheung
commented
Sep 4, 2019
is there a separate jira on this? |
HeartSaVioR
commented
Sep 4, 2019
Yes, filed https://issues.apache.org/jira/browse/SPARK-28967 and submitted a patch #25672 |
4bb9de0 to
082cf16CompareSparkQA
commented
Sep 4, 2019
Test build #110141 has finished for PR 25670 at commit
|
SparkQA
commented
Sep 5, 2019
Test build #110157 has finished for PR 25670 at commit
|
72a6253 to
9b4d53dCompareSparkQA
commented
Sep 5, 2019
Test build #110202 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 5, 2019
retest this, please |
SparkQA
commented
Sep 6, 2019
Test build #110208 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 6, 2019
First failure: known flaky test, SPARK-26989 (submitted a patch #25706) Second failure: SparkContext is leaked in test and makes tests failing. One of stack trace follows: |
HeartSaVioR
commented
Sep 6, 2019
retest this, please |
SparkQA
commented
Sep 6, 2019
Test build #110220 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 16, 2019
I'd like to put some efforts if that accelerates reviewing. Would it help reviewing if I split this to two parts - writer (EventLoggingListener) / reader (SHS)? |
HeartSaVioR
commented
Sep 16, 2019
Btw, I'll start working on next stuff which doesn't depend on this patch. Maybe I'll split them down into smaller parts than what I planned. Some parts may not be unused until following part will come. |
HeartSaVioR
commented
Sep 17, 2019
retest this, please |
SparkQA
commented
Sep 17, 2019
Test build #110704 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 17, 2019
UT failure: SPARK-29104 - not relevant. |
HeartSaVioR
commented
Sep 17, 2019
retest this, please |
HeartSaVioR
commented
Sep 17, 2019
FYI, #25811 is submitted to cover supporting snapshot/restore KVStore. |
SparkQA
commented
Sep 17, 2019
Test build #110730 has finished for PR 25670 at commit
|
HeartSaVioR
commented
Sep 17, 2019
retest this, please |
SparkQA
commented
Sep 17, 2019
Test build #110733 has finished for PR 25670 at commit
|
vanzin
commented
Oct 16, 2019
Looks ok, will merge tomorrow if no one else comments. |
SparkQA
commented
Oct 16, 2019
Test build #112190 has finished for PR 25670 at commit
|
vanzin
commented
Oct 17, 2019
Merging to master. |
HeartSaVioR
commented
Oct 17, 2019
Thanks all for reviewing and merging! |
gaborgsomogyi
commented
Nov 6, 2019
@HeartSaVioR just gone through this PR and plan to join later developments. You can ping me if this feature goes forward. One question what I have now. Do I see correctly that the actual implementation measures the events size before compression? If yes maybe my suggestion can be considered. Namely I can see 2 possibilities to overcame this but both cases the basic idea is the same. Dstream variable can be wrapped with
I've tested the second approach with lz4, lzf, snappy, zstd and only lz4 didn't flush the buffer immediately. Of course this doesn't mean 2nd approach is advised, just wanted to give more info... |
| private[spark] val EVENT_LOG_ROLLING_MAX_FILE_SIZE = | ||
| ConfigBuilder("spark.eventLog.rolling.maxFileSize") | ||
| .doc("The max size of event log file to be rolled over.") |
There was a problem hiding this comment.
Sorry leaving a comment late like this but it should have been better to say this configuration is only effective when spark.eventLog.rolling.enabled is enabled.
There was a problem hiding this comment.
I think there're counter examples in Spark configurations which rely on the fact once there's a configuration .enabled, others are effective only when that is enabled.
Even we only check from the SHS configuration, spark.history.fs.cleaner.*, spark.history.kerberos.*, spark.history.ui.acls.* fall into the case.
There was a problem hiding this comment.
Hm, I tend to disagree with omitting such dependent configurations in their documentations. Can we add and link related configurations in the documentations?
There was a problem hiding this comment.
Sorry. Looks like we'll have to agree to disagree then. No one has privilege to make someone do the work under his/her authorship which he/she disagrees with - it will end up putting wrong authorship on commit.
There was a problem hiding this comment.
@HeartSaVioR, it had to be reviewed. I just happened to review and leave some comments late. Logically if that's not documented, how do users know what configuration is effective when? At least I had to read the codes to confirm.
Also, I am trying to make sure we're on the same page so I wouldn't happen to leave this comment again since you are a regular contributor. I don't think this is a good pattern to don't document the relationship between configurations. I am going to send an email to the dev list.
There was a problem hiding this comment.
@HyukjinKwon
Thanks for initiating the thread in dev mailing list. I'm following up the thread and will be back once we get some sort of consensus.
… for rolling event log ### What changes were proposed in this pull request? This patch addresses the post-hoc review comment linked here - #25670 (comment) ### Why are the changes needed? We would like to explicitly document the direct relationship before we finish up structuring of configurations. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? N/A Closes#27576 from HeartSaVioR/SPARK-28869-FOLLOWUP-doc. Authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>
… for rolling event log ### What changes were proposed in this pull request? This patch addresses the post-hoc review comment linked here - #25670 (comment) ### Why are the changes needed? We would like to explicitly document the direct relationship before we finish up structuring of configurations. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? N/A Closes#27576 from HeartSaVioR/SPARK-28869-FOLLOWUP-doc. Authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org> (cherry picked from commit 446b2d2) Signed-off-by: HyukjinKwon <gurwls223@apache.org>
… for rolling event log ### What changes were proposed in this pull request? This patch addresses the post-hoc review comment linked here - apache#25670 (comment) ### Why are the changes needed? We would like to explicitly document the direct relationship before we finish up structuring of configurations. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? N/A Closesapache#27576 from HeartSaVioR/SPARK-28869-FOLLOWUP-doc. Authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>
This patch is a part of [SPARK-28594](https://issues.apache.org/jira/browse/SPARK-28594) and design doc for SPARK-28594 is linked here: https://docs.google.com/document/d/12bdCC4nA58uveRxpeo8k7kGOI2NRTXmXyBOweSi4YcY/edit?usp=sharing This patch proposes adding new feature to event logging, rolling event log files via configured file size. Previously event logging is done with single file and related codebase (`EventLoggingListener`/`FsHistoryProvider`) is tightly coupled with it. This patch adds layer on both reader (`EventLogFileReader`) and writer (`EventLogFileWriter`) to decouple implementation details between "handling events" and "how to read/write events from/to file". This patch adds two properties, `spark.eventLog.rollLog` and `spark.eventLog.rollLog.maxFileSize` which provides configurable behavior of rolling log. The feature is disabled by default, as we only expect huge event log for huge/long-running application. For other cases single event log file would be sufficient and still simpler. This is a part of SPARK-28594 which addresses event log growing infinitely for long-running application. This patch itself also provides some option for the situation where event log file gets huge and consume their storage. End users may give up replaying their events and want to delete the event log file, but given application is still running and writing the file, it's not safe to delete the file. End users will be able to delete some of old files after applying rolling over event log. No, as the new feature is turned off by default. Added unit tests, as well as basic manual tests. Basic manual tests - ran SHS, ran structured streaming query with roll event log enabled, verified split files are generated as well as SHS can load these files, with handling app status as incomplete/complete. Closesapache#25670 from HeartSaVioR/SPARK-28869. Lead-authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Co-authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan@gmail.com> Signed-off-by: Marcelo Vanzin <vanzin@cloudera.com>
### What changes were proposed in this pull request? This PR aims to enable `spark.eventLog.rolling.enabled` by default for Apache Spark 4.0.0. ### Why are the changes needed? Since Apache Spark 3.0.0, we have been using event log rolling not only for **long-running jobs**, but also for **some failed jobs** to archive the partial event logs incrementally. - #25670 ### Does this PR introduce _any_ user-facing change? - No because `spark.eventLog.enabled` is disabled by default. - For the users with `spark.eventLog.enabled=true`, yes, `spark-events` directory will have different layouts. However, all 3.3+ `Spark History Server` can read both old and new event logs. I believe that the event log users are already using this configuration to avoid the loss of event logs for long-running jobs and some failed jobs. ### How was this patch tested? Pass the CIs. ### Was this patch authored or co-authored using generative AI tooling? No. Closes#43638 from dongjoon-hyun/SPARK-45771. Authored-by: Dongjoon Hyun <dhyun@apple.com> Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
### What changes were proposed in this pull request? This PR aims to enable `spark.eventLog.rolling.enabled` by default for Apache Spark 4.0.0. ### Why are the changes needed? Since Apache Spark 3.0.0, we have been using event log rolling not only for **long-running jobs**, but also for **some failed jobs** to archive the partial event logs incrementally. - apache#25670 ### Does this PR introduce _any_ user-facing change? - No because `spark.eventLog.enabled` is disabled by default. - For the users with `spark.eventLog.enabled=true`, yes, `spark-events` directory will have different layouts. However, all 3.3+ `Spark History Server` can read both old and new event logs. I believe that the event log users are already using this configuration to avoid the loss of event logs for long-running jobs and some failed jobs. ### How was this patch tested? Pass the CIs. ### Was this patch authored or co-authored using generative AI tooling? No. Closesapache#43638 from dongjoon-hyun/SPARK-45771. Authored-by: Dongjoon Hyun <dhyun@apple.com> Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
…g.maxFileSize` ### What changes were proposed in this pull request? This PR aims to lower the minimum limit of `spark.eventLog.rolling.maxFileSize` from `10MiB` to `2MiB` at Apache Spark 4.1.0 while keeping the default (128MiB). ### Why are the changes needed? `spark.eventLog.rolling.maxFileSize` has `10MiB` as the lower bound limit since Apache Spark 3.0.0. - #25670 By reducing the lower bound to `2MiB`, we can allow Spark jobs to write small log files more frequently and faster without waiting for `10MiB`. This is helpful some slow(large micro-batch period) or low-traffic streaming jobs. The users will set a proper value for their jobs. ### Does this PR introduce _any_ user-facing change? There is no behavior change for the existing jobs. This only extends the range of configuration values for a user who wants to have lower values. ### How was this patch tested? Pass the CIs. ### Was this patch authored or co-authored using generative AI tooling? No. Closes#51162 from dongjoon-hyun/SPARK-52456. Authored-by: Dongjoon Hyun <dongjoon@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>
…in `History Server`
### What changes were proposed in this pull request?
This PR aims to support `On-Demand Log Loading` in `History Server` by looking up the **rolling event log locations** even Spark listing didn't finish to load the event log files.
```scala
val EVENT_LOG_ROLLING_ON_DEMAND_LOAD_ENABLED =
ConfigBuilder("spark.history.fs.eventLog.rolling.onDemandLoadEnabled")
.doc("Whether to look up rolling event log locations on demand manner before listing files.")
.version("4.1.0")
.booleanConf
.createWithDefault(true)
```
Previously, Spark History Server will show `Application ... Not Found` page if a job is requested before scanning it even if the file exists in the correct location. So, this PR doesn't introduce any regressions because this aims to introduce a kind of fallback logic to improve error handling .
<img width="686" height="359" alt="Screenshot 2025-07-22 at 14 08 21" src="https://github.com/user-attachments/assets/fccb413c-5a57-4918-86c0-28ae81d54873" />
### Why are the changes needed?
Since Apache Spark 3.0, we have been using event log rolling not only for **long-running jobs**, but also for **some failed jobs** to archive the partial event logs incrementally.
- #25670
Since Apache Spark 4.0, event log rolling is enabled by default.
- #43638
On top of that, this PR aims to reduce storage cost at Apache Spark 4.1. By supporting `On-Demand Loading for rolled event logs`, we can use larger values for `spark.history.fs.update.interval` instead of the default `10s`. Although Spark History logs are consumed in various ways, It has a big benefit because most of successful Spark jobs's logs are not visited by the users.
### Does this PR introduce _any_ user-facing change?
No. This is a new feature.
### How was this patch tested?
Pass the CIs with newly added test case.
### Was this patch authored or co-authored using generative AI tooling?
No.
Closes#51604 from dongjoon-hyun/SPARK-52914.
Authored-by: Dongjoon Hyun <dongjoon@apache.org>
Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>
What changes were proposed in this pull request?
This patch is a part of SPARK-28594 and design doc for SPARK-28594 is linked here: https://docs.google.com/document/d/12bdCC4nA58uveRxpeo8k7kGOI2NRTXmXyBOweSi4YcY/edit?usp=sharing
This patch proposes adding new feature to event logging, rolling event log files via configured file size.
Previously event logging is done with single file and related codebase (
EventLoggingListener/FsHistoryProvider) is tightly coupled with it. This patch adds layer on both reader (EventLogFileReader) and writer (EventLogFileWriter) to decouple implementation details between "handling events" and "how to read/write events from/to file".This patch adds two properties,
spark.eventLog.rollLogandspark.eventLog.rollLog.maxFileSizewhich provides configurable behavior of rolling log. The feature is disabled by default, as we only expect huge event log for huge/long-running application. For other cases single event log file would be sufficient and still simpler.Why are the changes needed?
This is a part of SPARK-28594 which addresses event log growing infinitely for long-running application.
This patch itself also provides some option for the situation where event log file gets huge and consume their storage. End users may give up replaying their events and want to delete the event log file, but given application is still running and writing the file, it's not safe to delete the file. End users will be able to delete some of old files after applying rolling over event log.
Does this PR introduce any user-facing change?
No, as the new feature is turned off by default.
How was this patch tested?
Added unit tests, as well as basic manual tests.
Basic manual tests - ran SHS, ran structured streaming query with roll event log enabled, verified split files are generated as well as SHS can load these files, with handling app status as incomplete/complete.