feat(cluster): 停用改为「标记」语义,让 disabled 真正随令牌环跨节点传播

承接用户提问「设计上停用不是本来就会跨节点传输吗」——核实结论:结构上确实
如此(TopoEntry.Link 是完整 store.Link,整个 State 随 token 每轮广播),但
实际路径断了。断点正是「撤销会删掉 topology 条目」:条目是 flag 的载体,
删了就无处传播,于是停用只能靠一次性 revoke 任务投递给 owner,**owner 当时
不在线就收不到**(实测 .60 记 disabled=1 / .106 记 0,就是这么来的)。

## 改为标记而非移除

撤销不再 RemoveTopology,而是 UpdateTopologyDisabled(true),条目保留、
Link.Disabled=true、Active=false。Active 正是为此存在:OfflineReassign()
只处理 Active 条目,所以停用的转发在 owner 掉线时不会被重新排队。

- 新增 UpdateTopologyDisabled / TopologyDisabled(照 UpdateTopologyGroup 的桥)
- 新增 store.ReconcileLinkDisabled 作接收端:adoption 时把环上的 flag 落进
  本地 store;本节点没有该转发时补一条 disabled 占位行(否则日后在本节点被
  claim 会复活),enable 则不建行
- SetTopologySync 由单向(store→环)扩为双向:群组仍上行,disabled 下行
- AddTopology 的 Active 跟随 Link.Disabled(原本硬编码 true,认领一个停用
  转发就会复活它)
- 审计日志细分 forward.stop / forward.start,与 forward.remove 区分

## 语义变更带出的两个新问题(都已修)

1. **「启动」这条路断了**。条目保留 ⇒ SubmitTask 被去重挡下,而认领路径的
   duplicate-claim 防御又会丢弃「已有 owner」的任务 ⇒ 重启任务发不出去,owner
   永远收不到,转发**能停不能起**。
   修:新增 Task.Restart 这一独立任务类型 + SubmitRestart + Handler.RestartFn,
   显式绕过 duplicate-claim 防御并原地复活(不重复建条目、不重跑 claim 簿记)。
   SubmitTask 的守卫同时从 HasTask 收窄为新的 HasActiveTask(跳过 disabled 条目
   与撤销任务);saveCanvas 的判断相应改用 HasActiveTask,避免每次保存都对
   已标记的转发重复发撤销。

2. 原本两处 RemoveTopologyEntry 调用(ClaimFn/RevokeFn 的 disabled 分支)在
   新语义下会把本该保留的条目删掉,改为 UpdateTopologyDisabled。

## 测试(每个都做了「回退修复行→必须变红→还原变绿」双向验证)

- TestStoppedTopologyEntrySurvivesAdoption —— 离线成员也能学到停用,
  一次性 revoke 任务永远做不到这一点
- TestStoppedForwardNotRequeuedOnNodeDeparture / TestAddTopologyRespectsDisabledFlag
  —— 标记而非删除为何安全
- TestSubmitTaskNotBlockedByStoppedEntry / TestSubmitTaskStillDedupesActiveForward
- TestRestartTaskBypassesDuplicateClaimGuard / TestRestartFlagSurvivesTokenSerialization
- TestStopThenStartPublishesRestartTask(HTTP 端到端,断言**任务真的发出**)
- TestReconcileLinkDisabled*(store 侧三条)

★ 两次踩到**假绿**:第一版只断言 store 层(newTestHandler 的 Ring 为 nil,
坏掉的路根本没执行);第二版在 re-enable **之后**才调 SubmitTask,此时新旧
谓词结果相同,测不出差异。都是靠「回退修复行看是否变红」抓出来的 —— 这个
双向验证已经是本项目的固定动作。

go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
JianFeeeee
2026-09-26 10:44:04 +08:00
parent 041cc04dd6
commit 1c835425de
12 changed files with 830 additions and 48 deletions

View File

@ -62,6 +62,12 @@ type Task struct {
// owning node cancels it (stop worker, drop from topology). Reuses the
// same publish channel as creation (round-1 inject, round-2 apply).
Revoke bool `json:"revoke,omitempty"`
// Restart marks a RE-ENABLE task: the forward is already claimed and its
// topology entry still exists (a stop keeps the entry as the flag's
// carrier), so the creation path would dedupe the submission and the
// duplicate-claim guard would discard it. The owner applies this one
// unconditionally: re-mark the entry enabled and (re)spawn the worker.
Restart bool `json:"restart,omitempty"`
// RemoveNode: node ID to remove from the ring; the node self-removes when
// the command reaches it via the token.
RemoveNode string `json:"removeNode,omitempty"`
@ -333,10 +339,15 @@ func (s *State) AddTopology(t *Task, ownerID string) *TopoEntry {
if s.Topology == nil {
s.Topology = map[string]*TopoEntry{}
}
// Active mirrors the task's disabled flag rather than being unconditionally
// true. A claim can legitimately carry a stopped forward (e.g. a node
// re-claiming from a departed peer), and forcing Active=true there would
// resurrect it: OfflineReassign only re-queues Active entries, and every
// reader that filters on Active would treat it as running again.
e := &TopoEntry{
TaskID: t.ID, OwnerID: ownerID,
Local: t.Local, Remote: t.Remote, Link: t.Link,
Active: true,
Active: !t.Link.Disabled,
}
s.Topology[t.ID] = e
return e
@ -408,6 +419,39 @@ func (s *State) ForwardsOwnedBy(owner string) []*TopoEntry {
// (local, remote, port) triple. Returns true if found. The updated entry
// propagates to all nodes via the next token cycle — group changes sync
// through the ring without a dedicated command.
// UpdateTopologyDisabled sets the disabled flag on the topology entry matching
// the forward's natural key. Mirror of UpdateTopologyGroup: it makes a locally
// decided stop/start part of the topology so the next token cycle carries it to
// every other member (the alternative — a one-shot revoke task addressed at the
// owner — only converges if that owner happens to be online right then).
//
// Returns false when the forward is not in the topology, which is not an error:
// a stopped forward may legitimately have no entry yet.
func (s *State) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {
e.Link.Disabled = disabled
// Keep Active consistent with the flag so readers that key off it
// (status page, load accounting) agree with Link.Disabled.
e.Active = !disabled
return true
}
}
return false
}
// TopologyDisabled returns the disabled flag the cluster currently holds for a
// forward, and whether such an entry exists. Used by the adoption path to learn
// a peer's decision into the local store.
func (s *State) TopologyDisabled(local, remote string, port int) (bool, bool) {
for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {
return e.Link.Disabled, true
}
}
return false, false
}
func (s *State) UpdateTopologyGroup(local, remote string, port int, group string) bool {
for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {