fix(cluster): 停用状态以「环」为权威,消除长期离线导致的永久分歧

承接用户指出的遗留:links 表是每节点本地副本,靠「store→环上行 + 环→store
下行」双向 sync 收敛,但两边都可能被覆盖。其中「本地 store 权威」这条规则有
硬伤,本次改掉。

## 先复现,再动手

写探针验证「入环 token 会不会冲掉本地未传播的决定」,结果比预期严重:

    after adopting a stale token: known=true disabled=false
    LOCAL DISABLE WAS WIPED by an incoming token

OnToken/AdoptState 是 `e.state = tk.State` **整体替换**。两次 token 之间做出的
停用决定,只要下一轮到达的 token 是「决定之前」捕获的,就会被整个洗掉 ——
决定永远传不出去,转发照旧运行。Group 之所以看起来没这个问题,是因为它每次
adoption 都被 `SetTopologySync` 从 store 重新推上去,而 disabled 没有对应的
「尚未传播」保护。

## 改法:环权威 + 本地决定带「未确认」标记

1. **环权威**:adoption 时把环上的 disabled 写穿本地 store(write-through,
   不是监听器),任何 peer 的决定都在一个 token 周期内落地。这终结了旧规则
   「各人信自己那份」造成的永久分歧。顺序上只有单向要求:引擎先重新断言本地
   未确认决定,再跑 host sync,所以读环绝不会覆盖用户刚做的操作。

2. **未确认决定受保护**(RingEngine.localDisabled):`UpdateTopologyDisabled`
   记下决定,adoption 后由 `reconcileLocalDisabled` 重新断言到刚采纳的 state 上,
   于是它会随下一轮 token 传出去。环报回同值时删除该键(全cluster已一致);
   转发从 topology 消失时一并清扫,map 不会无限增长。**false 同样受保护** ——
   重新启用也需要传播,丢掉它会把转发永久留在停用态。

3. `TopologyDisabled()` 对未确认决定短路返回本地值,避免状态页与用户刚做的
   操作相反。

## 测试(又抓到一个假绿)

新增 4 条:停用/启用跨 adoption 存活、确认后停止断言(让位给 peer 的后续决定)、
追踪表不累积。

★ `TestLocalDisableSurvivesAdoption` **第一版是假绿**:它断言
`TopologyDisabled()`,而该访问器会短路到本地决定,于是「即便即将转发的 state
仍是 enabled」它也报 true。改成断言**下游节点会看到什么**(用一个 peer 引擎
AdoptState 本引擎的 state)后,去掉重新断言如期变红:

    the forwarded token carries disabled=false, want true

双向验证通过,这个教训要记:测「声明」而不是测「实际传播的状态」,等于没测。

go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
JianFeeeee
2026-09-26 11:44:28 +08:00
parent c8ab2584a8
commit 918ca5d5ba
3 changed files with 268 additions and 15 deletions

View File

@ -7,6 +7,7 @@ import (
"encoding/json"
"fmt"
"log"
"strconv"
"sync"
"time"
@ -55,6 +56,24 @@ type Engine struct {
// group changes made via HTTP handlers between token cycles are overwritten
// by the next state adoption and never propagate to other nodes.
topologySync func()
// localDisabled records this node's own enabled/disabled decisions that have
// not yet been confirmed by the ring. It is what makes the RING authoritative
// without losing a decision that was made locally and has not yet had a
// chance to reach the other members.
//
// The problem it solves: OnToken/AdoptState do `e.state = tk.State`, a
// wholesale replacement. A stop decided between two token cycles therefore
// disappears on the very next adoption if that token was captured before the
// decision — the flag never reaches the rest of the ring, and the forward
// keeps running. (Verified with a probe before writing this: adopting a
// stale token flipped disabled back to false.)
//
// A key stays in this map until an adoption reports it disabled, at which
// point the whole cluster agrees and the key is dropped. Keys for forwards
// that vanish from the topology are swept on every update, so the map cannot
// grow without bound. The value false is what keeps a RE-ENABLE alive for the
// same reason a stop needs protecting.
localDisabled map[string]bool
state State
// myAddr maps our Node ID to the address peers dial.
@ -236,6 +255,11 @@ func (e *Engine) OnToken(ctx context.Context, tk *Token) (*Token, error) {
// Re-apply local store overrides (group labels, disabled flags) onto
// the freshly adopted topology so they survive state adoption and
// propagate to all nodes via the next token forward.
// Re-assert decisions this node made locally that the ring has not yet
// confirmed, BEFORE the host's topology sync runs: the sync reads
// TopologyDisabled to learn the cluster's view, so the local decision must
// already be applied or a stop could be reported as "enabled" and dropped.
e.reconcileLocalDisabled()
if e.topologySync != nil {
e.topologySync()
}
@ -802,6 +826,11 @@ func (e *Engine) AdoptState(s State) {
for id, t := range kept {
e.state.PendingTasks[id] = t
}
// Re-assert decisions this node made locally that the ring has not yet
// confirmed, BEFORE the host's topology sync runs: the sync reads
// TopologyDisabled to learn the cluster's view, so the local decision must
// already be applied or a stop could be reported as "enabled" and dropped.
e.reconcileLocalDisabled()
if e.topologySync != nil {
e.topologySync()
}
@ -990,16 +1019,75 @@ func (e *Engine) UpdateTopologyGroup(local, remote string, port int, group strin
// UpdateTopologyDisabled marks a forward enabled/disabled in the topology so
// the decision rides the next token cycle to every member. Paired with
// TopologyDisabled, which the adoption hook uses to learn peers' decisions.
//
// The decision is also remembered locally until the ring confirms it, because
// adoption replaces the whole state: a token captured before this call would
// otherwise wash the decision out on the next cycle and it would never reach
// the other members. See the localDisabled field.
func (e *Engine) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
if e.localDisabled == nil {
e.localDisabled = map[string]bool{}
}
e.localDisabled[disableKey(local, remote, port)] = disabled
return e.state.UpdateTopologyDisabled(local, remote, port, disabled)
}
// TopologyDisabled reports the cluster's view of a forward's disabled flag.
// The bool is false when the forward has no topology entry (nothing to learn).
//
// A pending local decision takes precedence over the adopted state: it has not
// had a chance to reach the other members yet, and reporting the adopted value
// would make the local store (and the status page) disagree with the user's
// most recent action.
func (e *Engine) TopologyDisabled(local, remote string, port int) (bool, bool) {
if d, pending := e.localDisabled[disableKey(local, remote, port)]; pending {
return d, true
}
return e.state.TopologyDisabled(local, remote, port)
}
// reconcileLocalDisabled re-asserts this node's not-yet-confirmed disabled
// decisions onto the freshly adopted state. Called right after e.state is
// replaced, and paired with reconcileTopologySync (which pushes those decisions
// out to the ring).
//
// A decision is dropped once the adopted state reports the SAME value: at that
// point every member agrees and there is nothing left to protect. Entries whose
// forward no longer exists in the topology are dropped too, so the map tracks
// only live disagreements.
func (e *Engine) reconcileLocalDisabled() {
if len(e.localDisabled) == 0 {
return
}
for id, te := range e.state.Topology {
_ = id
k := disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)
want, pending := e.localDisabled[k]
if !pending {
continue
}
if te.Link.Disabled == want {
// The ring caught up (or agreed independently): stop tracking.
delete(e.localDisabled, k)
continue
}
// Still divergent: keep asserting our decision onto the adopted state.
te.Link.Disabled = want
te.Active = !want
}
// Forget decisions for forwards that left the topology entirely — otherwise
// a long-lived cluster would accumulate dead keys.
live := make(map[string]struct{}, len(e.state.Topology))
for _, te := range e.state.Topology {
live[disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)] = struct{}{}
}
for k := range e.localDisabled {
if _, ok := live[k]; !ok {
delete(e.localDisabled, k)
}
}
}
// IsLeader reports whether this node is the current ring leader.
func (e *Engine) IsLeader() bool { return e.state.LeaderID == e.ID }
@ -1022,6 +1110,14 @@ func (e *Engine) SetPeerPersist(fn func(peersJSON string) error) { e.peerPersist
// HTTP handlers are overwritten by the next e.state = tk.State.
func (e *Engine) SetTopologySync(fn func()) { e.topologySync = fn }
// disableKey builds the map key for a forward's disabled decision: the natural
// triple (local, remote, remotePort), which is how every other part of the code
// identifies a forward. A NUL separator keeps it unambiguous for names that
// could otherwise collide across the boundaries.
func disableKey(local, remote string, port int) string {
return local + "\x00" + remote + "\x00" + strconv.Itoa(port)
}
// persistPeers extracts all alive peers (addr + nodeKey, excluding self)
// from the current ring state and persists them via the peerPersist callback.
// Called on every token cycle (OnToken) and on AdoptState so a crashed node