Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); RATIS-1866. Maintain leader lease after AppendEntries by SzyWilliam · Pull Request #898 · apache/ratis · GitHub
Skip to content

RATIS-1866. Maintain leader lease after AppendEntries - #898

Merged
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866
Sep 22, 2023
Merged

RATIS-1866. Maintain leader lease after AppendEntries#898
SzyWilliam merged 5 commits into
apache:feature/leaderleasefrom
SzyWilliam:jira1866

Conversation

@SzyWilliam

@SzyWilliamSzyWilliam commented Aug 4, 2023

Copy link
Copy Markdown
Member

What is a Leader Lease

In Raft, the leader is responsible for processing and coordinating client requests, replicating data among other followers, and maintaining the distributed state machine.
Vanilla Raft requires the leader to obtain majority acknowledgements before serving every read requests. During normal operations, this prerequisite leads to unnecessary rpcs exchanged among the cluster, diminished read throughput and increased latency.
The leader lease is a concept that allows the leader to maintain its leadership without obtaining majority acknowledgements for a certain period of time (lease duration), during which it can directly serve client read requests.

How to extend lease during normal operations

Prerequisite

  • Suppose the CPU clocks are perfectly synchronized among cluster machines.
  • Let each AppendEntries request AE(i) contain a send time T(i)

Initialize

Once a leader is elected and its authority being comfirmed by majorities through successfully replicating its first no-op log, the leader gains the lease. The lease validity starts from T(0).

Renewal

As long as the leader continues to send heartbeats and receives acknowledgments from a majority of other nodes, it can renew its lease. Theoretically, if the most recent acknowledged heartbeat was sent at time T(n), the validity of the new lease commences at T(n).
In practice, rather than updating the lease with every heartbeat, we opt for a more efficient approach by lazily updating the leader's lease upon each query. Here's how it works:
At time T(n), when the leader is questioned about its authority, it first collects the send times of the last replied AppendEntries from each of its followers, denoted as TR(1), TR(2), ..., TR(2n), sorted in descending order.
Next, it selects the maximum timestamp at when the majority of followers are known to be active, that is, TR(n).
If TR(n) falls within the time range [T(n), T(n) + LeaseTimeoutDuration], then the lease can be successfully renewed.

Revoke

If the lease is expired and the leader cannot renew it, it loses the lease and stops serving read-only requests directly.

How to handle lease during configuration changes

During the configuration changes, the lease can only be renewed if acknowledgments be received by both the old group and the new group. It is the same to leader election restrictions during reconfiguration.

What to do when forced step down

When a leader is forced down, its lease should be effectively revoked.

How to handle CPU drifts

We can lower the ratio allowed for lease timeouts. If the CPU drifts are unbound, better not to use lease read :)

See https://issues.apache.org/jira/browse/RATIS-1866.

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! Please see the comments inlined.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for this detailed review! I will address these issues a bit later ;)

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks a lot for working on this! I have some questions and comments inlined. The changing leader case is tricky.

Comment threadratis-grpc/src/main/java/org/apache/ratis/grpc/server/GrpcLogAppender.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
Comment on lines +252 to +253
return Stream.concat(current.stream(),
Optional.ofNullable(old).map(List::stream).orElse(Stream.empty()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should deduplicate the peers.

Comment threadratis-server/src/main/java/org/apache/ratis/server/impl/LeaderLease.java Outdated
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo Thanks very much for the detailed review! I'll elaborate the leader changing process. (may be in next PR).


class LeaderLease {

private final long leaseTimeoutMs;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make it static?

getFollower().updateLastRpcSendTime(request.getEntriesCount() == 0);
final AppendEntriesReplyProto r = getServerRpc().appendEntries(request);
getFollower().updateLastRpcResponseTime();
getFollower().updateLastRespondedAppendEntriesSendTime(sendTime);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why still update LastRespondedAppendEntriesSendTime at the time of sending?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getServerRpc().appendEntries(request) is a blocking operation, and once this call returns, the response for the current AppendEntries request has been received. Therefore, we can update its(LastRespondedAppendEntries)sendTime.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it~

@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

Made changes on code. @szetszwo@OneSizeFitsQuorum PTAL, thanks!

@OneSizeFitsQuorumOneSizeFitsQuorum left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SzyWilliam , thanks for the update! The change looks good. Just a comment inlined.

Comment on lines +80 to +91
private Timestamp getMaxTimestampWithMajorityAck(List<FollowerInfo> peers) {
if (peers == null || peers.isEmpty()) {
return Timestamp.currentTime();
}

final List<Timestamp> lastRespondedAppendEntriesSendTimes = peers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.collect(Collectors.toList());

return lastRespondedAppendEntriesSendTimes.get(lastRespondedAppendEntriesSendTimes.size() / 2);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the leader is not in the peer list, we should use (lastRespondedAppendEntriesSendTimes.size() - 1)/ 2:

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Instead of creating a list, we may use limit and skip as below.

privateTimestampgetMaxTimestampWithMajorityAck(List<FollowerInfo> followers) {
if (followers == null || followers.isEmpty()) {
returnTimestamp.currentTime();
}
finalintmid = (followers.size() - 1) / 2;
returnfollowers.stream()
.map(FollowerInfo::getLastRespondedAppendEntriesSendTime)
.sorted()
.limit(mid + 1)
.skip(mid)
.iterator()
.next();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • 1 or 2 followers: use index 0
  • 3 or 4 followers: use index 1

Oops, the timestamps are sorted in ascending order but not descending order. Then it should be

  • 1 follower: use index 0
  • 2 or 3 followers: use index 1
  • 4 or 5 followers: use index 2

You formula actually is correct!

finalintmid = followers.size() / 2;

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot for the reviews! Didn't know we can use limit and skip. Now the code is more light-weighted!

@szetszwoszetszwo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 the change looks good.

@SzyWilliam
SzyWilliam merged commit a13d81e into apache:feature/leaderleaseSep 22, 2023
@SzyWilliam

Copy link
Copy Markdown
MemberAuthor

@szetszwo@OneSizeFitsQuorum Thanks a lot for your careful and thorough reviews!

RexXiong pushed a commit to apache/celeborn that referenced this pull request May 30, 2024
### What changes were proposed in this pull request?
Bump Ratis version from 2.5.1 to 3.0.1. Address incompatible changes:
- RATIS-589. Eliminate buffer copying in SegmentedRaftLogOutputStream.(apache/ratis#964)
- RATIS-1677. Do not auto format RaftStorage in RECOVER.(apache/ratis#718)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
### Why are the changes needed?
Bump Ratis version from 2.5.1 to 3.0.1. Ratis has released v3.0.0, v3.0.1, which release note refers to [3.0.0](https://ratis.apache.org/post/3.0.0.html), [3.0.1](https://ratis.apache.org/post/3.0.1.html). The 3.0.x version include new features like pluggable metrics and lease read, etc, some improvements and bugfixes including:
- 3.0.0: Change list of ratis 3.0.0 In total, there are roughly 100 commits diffing from 2.5.1 including:
- Incompatible Changes
- RaftStorage Auto-Format
- RATIS-1677. Do not auto format RaftStorage in RECOVER. (apache/ratis#718)
- RATIS-1694. Fix the compatibility issue of RATIS-1677. (apache/ratis#731)
- RATIS-1871. Auto format RaftStorage when there is only one directory configured. (apache/ratis#903)
- Pluggable Ratis-Metrics (RATIS-1688)
- RATIS-1689. Remove the use of the thirdparty Gauge. (apache/ratis#728)
- RATIS-1692. Remove the use of the thirdparty Counter. (apache/ratis#732)
- RATIS-1693. Remove the use of the thirdparty Timer. (apache/ratis#734)
- RATIS-1703. Move MetricsReporting and JvmMetrics to impl. (apache/ratis#741)
- RATIS-1704. Fix SuppressWarnings(“VisibilityModifier”) in RatisMetrics. (apache/ratis#742)
- RATIS-1710. Refactor metrics api and implementation to separated modules. (apache/ratis#749)
- RATIS-1712. Add a dropwizard 3 implementation of ratis-metrics-api. (apache/ratis#751)
- RATIS-1391. Update library dropwizard.metrics version to 4.x (apache/ratis#632)
- RATIS-1601. Use the shaded dropwizard metrics and remove the dependency (apache/ratis#671)
- Streaming Protocol Change
- RATIS-1569. Move the asyncRpcApi.sendForward(..) call to the client side. (apache/ratis#635)
- New Features
- Leader Lease (RATIS-1864)
- RATIS-1865. Add leader lease bound ratio configuration (apache/ratis#897)
- RATIS-1866. Maintain leader lease after AppendEntries (apache/ratis#898)
- RATIS-1894. Implement ReadOnly based on leader lease (apache/ratis#925)
- RATIS-1882. Support read-after-write consistency (apache/ratis#913)
- StateMachine API
- RATIS-1874. Add notifyLeaderReady function in IStateMachine (apache/ratis#906)
- RATIS-1897. Make TransactionContext available in DataApi.write(..). (apache/ratis#930)
- New Configuration Properties
- RATIS-1862. Add the parameter whether to take Snapshot when stopping to adapt to different services (apache/ratis#896)
- RATIS-1930. Add a conf for enable/disable majority-add. (apache/ratis#961)
- RATIS-1918. Introduces parameters that separately control the shutdown of RaftServerProxy by JVMPauseMonitor. (apache/ratis#950)
- RATIS-1636. Support re-config ratis properties (apache/ratis#800)
- RATIS-1860. Add ratis-shell cmd to generate a new raft-meta.conf. (apache/ratis#901)
- Improvements & Bug Fixes
- Netty
- RATIS-1898. Netty should use EpollEventLoopGroup by default (apache/ratis#931)
- RATIS-1899. Use EpollEventLoopGroup for Netty Proxies (apache/ratis#932)
- RATIS-1921. Shared worker group in WorkerGroupGetter should be closed. (apache/ratis#955)
- RATIS-1923. Netty: atomic operations require side-effect-free functions. (apache/ratis#956)
- RaftServer
- RATIS-1924. Increase the default of raft.server.log.segment.size.max. (apache/ratis#957)
- RATIS-1892. Unify the lifetime of the RaftServerProxy thread pool (apache/ratis#923)
- RATIS-1889. NoSuchMethodError: RaftServerMetricsImpl.addNumPendingRequestsGauge apache/ratis#922 (apache/ratis#922)
- RATIS-761. Handle writeStateMachineData failure in leader. (apache/ratis#927)
- RATIS-1902. The snapshot index is set incorrectly in InstallSnapshotReplyProto. (apache/ratis#933)
- RATIS-1912. Fix infinity election when perform membership change. (apache/ratis#954)
- RATIS-1858. Follower keeps logging first election timeout. (apache/ratis#894)
- 3.0.1:This is a bugfix release. See the [changes between 3.0.0 and 3.0.1](apache/ratis@ratis-3.0.0...ratis-3.0.1) releases.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Cluster manual test.
Closes#2480 from SteNicholas/CELEBORN-1400.
Authored-by: SteNicholas <programgeek@163.com>
Signed-off-by: Shuang <lvshuang.xjs@alibaba-inc.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SzyWilliam@szetszwo@OneSizeFitsQuorum