mirror of
https://gitcode.com/JianFeeeee/webui4frpc.git
synced 2026-10-03 15:43:59 +00:00
fix(cluster): 停用状态以「环」为权威,消除长期离线导致的永久分歧
承接用户指出的遗留:links 表是每节点本地副本,靠「store→环上行 + 环→store
下行」双向 sync 收敛,但两边都可能被覆盖。其中「本地 store 权威」这条规则有
硬伤,本次改掉。
## 先复现,再动手
写探针验证「入环 token 会不会冲掉本地未传播的决定」,结果比预期严重:
after adopting a stale token: known=true disabled=false
LOCAL DISABLE WAS WIPED by an incoming token
OnToken/AdoptState 是 `e.state = tk.State` **整体替换**。两次 token 之间做出的
停用决定,只要下一轮到达的 token 是「决定之前」捕获的,就会被整个洗掉 ——
决定永远传不出去,转发照旧运行。Group 之所以看起来没这个问题,是因为它每次
adoption 都被 `SetTopologySync` 从 store 重新推上去,而 disabled 没有对应的
「尚未传播」保护。
## 改法:环权威 + 本地决定带「未确认」标记
1. **环权威**:adoption 时把环上的 disabled 写穿本地 store(write-through,
不是监听器),任何 peer 的决定都在一个 token 周期内落地。这终结了旧规则
「各人信自己那份」造成的永久分歧。顺序上只有单向要求:引擎先重新断言本地
未确认决定,再跑 host sync,所以读环绝不会覆盖用户刚做的操作。
2. **未确认决定受保护**(RingEngine.localDisabled):`UpdateTopologyDisabled`
记下决定,adoption 后由 `reconcileLocalDisabled` 重新断言到刚采纳的 state 上,
于是它会随下一轮 token 传出去。环报回同值时删除该键(全cluster已一致);
转发从 topology 消失时一并清扫,map 不会无限增长。**false 同样受保护** ——
重新启用也需要传播,丢掉它会把转发永久留在停用态。
3. `TopologyDisabled()` 对未确认决定短路返回本地值,避免状态页与用户刚做的
操作相反。
## 测试(又抓到一个假绿)
新增 4 条:停用/启用跨 adoption 存活、确认后停止断言(让位给 peer 的后续决定)、
追踪表不累积。
★ `TestLocalDisableSurvivesAdoption` **第一版是假绿**:它断言
`TopologyDisabled()`,而该访问器会短路到本地决定,于是「即便即将转发的 state
仍是 enabled」它也报 true。改成断言**下游节点会看到什么**(用一个 peer 引擎
AdoptState 本引擎的 state)后,去掉重新断言如期变红:
the forwarded token carries disabled=false, want true
双向验证通过,这个教训要记:测「声明」而不是测「实际传播的状态」,等于没测。
go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
@ -7,6 +7,7 @@ import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log"
|
||||
"strconv"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
@ -55,6 +56,24 @@ type Engine struct {
|
||||
// group changes made via HTTP handlers between token cycles are overwritten
|
||||
// by the next state adoption and never propagate to other nodes.
|
||||
topologySync func()
|
||||
// localDisabled records this node's own enabled/disabled decisions that have
|
||||
// not yet been confirmed by the ring. It is what makes the RING authoritative
|
||||
// without losing a decision that was made locally and has not yet had a
|
||||
// chance to reach the other members.
|
||||
//
|
||||
// The problem it solves: OnToken/AdoptState do `e.state = tk.State`, a
|
||||
// wholesale replacement. A stop decided between two token cycles therefore
|
||||
// disappears on the very next adoption if that token was captured before the
|
||||
// decision — the flag never reaches the rest of the ring, and the forward
|
||||
// keeps running. (Verified with a probe before writing this: adopting a
|
||||
// stale token flipped disabled back to false.)
|
||||
//
|
||||
// A key stays in this map until an adoption reports it disabled, at which
|
||||
// point the whole cluster agrees and the key is dropped. Keys for forwards
|
||||
// that vanish from the topology are swept on every update, so the map cannot
|
||||
// grow without bound. The value false is what keeps a RE-ENABLE alive for the
|
||||
// same reason a stop needs protecting.
|
||||
localDisabled map[string]bool
|
||||
|
||||
state State
|
||||
// myAddr maps our Node ID to the address peers dial.
|
||||
@ -236,6 +255,11 @@ func (e *Engine) OnToken(ctx context.Context, tk *Token) (*Token, error) {
|
||||
// Re-apply local store overrides (group labels, disabled flags) onto
|
||||
// the freshly adopted topology so they survive state adoption and
|
||||
// propagate to all nodes via the next token forward.
|
||||
// Re-assert decisions this node made locally that the ring has not yet
|
||||
// confirmed, BEFORE the host's topology sync runs: the sync reads
|
||||
// TopologyDisabled to learn the cluster's view, so the local decision must
|
||||
// already be applied or a stop could be reported as "enabled" and dropped.
|
||||
e.reconcileLocalDisabled()
|
||||
if e.topologySync != nil {
|
||||
e.topologySync()
|
||||
}
|
||||
@ -802,6 +826,11 @@ func (e *Engine) AdoptState(s State) {
|
||||
for id, t := range kept {
|
||||
e.state.PendingTasks[id] = t
|
||||
}
|
||||
// Re-assert decisions this node made locally that the ring has not yet
|
||||
// confirmed, BEFORE the host's topology sync runs: the sync reads
|
||||
// TopologyDisabled to learn the cluster's view, so the local decision must
|
||||
// already be applied or a stop could be reported as "enabled" and dropped.
|
||||
e.reconcileLocalDisabled()
|
||||
if e.topologySync != nil {
|
||||
e.topologySync()
|
||||
}
|
||||
@ -990,16 +1019,75 @@ func (e *Engine) UpdateTopologyGroup(local, remote string, port int, group strin
|
||||
// UpdateTopologyDisabled marks a forward enabled/disabled in the topology so
|
||||
// the decision rides the next token cycle to every member. Paired with
|
||||
// TopologyDisabled, which the adoption hook uses to learn peers' decisions.
|
||||
//
|
||||
// The decision is also remembered locally until the ring confirms it, because
|
||||
// adoption replaces the whole state: a token captured before this call would
|
||||
// otherwise wash the decision out on the next cycle and it would never reach
|
||||
// the other members. See the localDisabled field.
|
||||
func (e *Engine) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
|
||||
if e.localDisabled == nil {
|
||||
e.localDisabled = map[string]bool{}
|
||||
}
|
||||
e.localDisabled[disableKey(local, remote, port)] = disabled
|
||||
return e.state.UpdateTopologyDisabled(local, remote, port, disabled)
|
||||
}
|
||||
|
||||
// TopologyDisabled reports the cluster's view of a forward's disabled flag.
|
||||
// The bool is false when the forward has no topology entry (nothing to learn).
|
||||
//
|
||||
// A pending local decision takes precedence over the adopted state: it has not
|
||||
// had a chance to reach the other members yet, and reporting the adopted value
|
||||
// would make the local store (and the status page) disagree with the user's
|
||||
// most recent action.
|
||||
func (e *Engine) TopologyDisabled(local, remote string, port int) (bool, bool) {
|
||||
if d, pending := e.localDisabled[disableKey(local, remote, port)]; pending {
|
||||
return d, true
|
||||
}
|
||||
return e.state.TopologyDisabled(local, remote, port)
|
||||
}
|
||||
|
||||
// reconcileLocalDisabled re-asserts this node's not-yet-confirmed disabled
|
||||
// decisions onto the freshly adopted state. Called right after e.state is
|
||||
// replaced, and paired with reconcileTopologySync (which pushes those decisions
|
||||
// out to the ring).
|
||||
//
|
||||
// A decision is dropped once the adopted state reports the SAME value: at that
|
||||
// point every member agrees and there is nothing left to protect. Entries whose
|
||||
// forward no longer exists in the topology are dropped too, so the map tracks
|
||||
// only live disagreements.
|
||||
func (e *Engine) reconcileLocalDisabled() {
|
||||
if len(e.localDisabled) == 0 {
|
||||
return
|
||||
}
|
||||
for id, te := range e.state.Topology {
|
||||
_ = id
|
||||
k := disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)
|
||||
want, pending := e.localDisabled[k]
|
||||
if !pending {
|
||||
continue
|
||||
}
|
||||
if te.Link.Disabled == want {
|
||||
// The ring caught up (or agreed independently): stop tracking.
|
||||
delete(e.localDisabled, k)
|
||||
continue
|
||||
}
|
||||
// Still divergent: keep asserting our decision onto the adopted state.
|
||||
te.Link.Disabled = want
|
||||
te.Active = !want
|
||||
}
|
||||
// Forget decisions for forwards that left the topology entirely — otherwise
|
||||
// a long-lived cluster would accumulate dead keys.
|
||||
live := make(map[string]struct{}, len(e.state.Topology))
|
||||
for _, te := range e.state.Topology {
|
||||
live[disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)] = struct{}{}
|
||||
}
|
||||
for k := range e.localDisabled {
|
||||
if _, ok := live[k]; !ok {
|
||||
delete(e.localDisabled, k)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// IsLeader reports whether this node is the current ring leader.
|
||||
func (e *Engine) IsLeader() bool { return e.state.LeaderID == e.ID }
|
||||
|
||||
@ -1022,6 +1110,14 @@ func (e *Engine) SetPeerPersist(fn func(peersJSON string) error) { e.peerPersist
|
||||
// HTTP handlers are overwritten by the next e.state = tk.State.
|
||||
func (e *Engine) SetTopologySync(fn func()) { e.topologySync = fn }
|
||||
|
||||
// disableKey builds the map key for a forward's disabled decision: the natural
|
||||
// triple (local, remote, remotePort), which is how every other part of the code
|
||||
// identifies a forward. A NUL separator keeps it unambiguous for names that
|
||||
// could otherwise collide across the boundaries.
|
||||
func disableKey(local, remote string, port int) string {
|
||||
return local + "\x00" + remote + "\x00" + strconv.Itoa(port)
|
||||
}
|
||||
|
||||
// persistPeers extracts all alive peers (addr + nodeKey, excluding self)
|
||||
// from the current ring state and persists them via the peerPersist callback.
|
||||
// Called on every token cycle (OnToken) and on AdoptState so a crashed node
|
||||
|
||||
Reference in New Issue
Block a user