mirror of
https://gitcode.com/JianFeeeee/webui4frpc.git
synced 2026-10-03 23:53:59 +00:00
fix(cluster): restart 任务必须定向投递给 owner,不能被非 owner 吞掉
上一提交(5cdc052)只加了执行侧 owner 判定,实测仍然失败:从 .60(非 owner)
启动 owner 在 .30 的 portal,.30 的 worker 一直没起来,且三台日志里既没有
`restarted` 也没有任何错误。
## 真因:任务被非 owner「消费」掉了
pending 命令由 runCommands 的 `for {}` 循环每轮重取 PendingList(),取出后
ClaimPending 即从 map 移除。我当时在「owner != e.ID」时把任务塞回
PendingTasks —— 于是它**立刻又变回待处理**,下一轮循环再次取到,无限
`defer restart`(单测直接跑成死循环,300s 超时)。
而上一版的 `continue` 同样是错的:ClaimPending 已经把任务移除,continue
等于消费,owner 永远收不到。
## 修法:放进「选择」循环,和 Revoke 完全同构
runCommands 里 Revoke 早就有正确的定向投递范式:
owner == e.ID → 我持有,执行
owner == "" 且最低负载 → 转发已消失,兜底消费
否则 → continue(任务**留在 token 里**随环前进)
restart 照抄这套。非 owner 只是不选中它,任务随 token 传给下一个节点,直到
owner 那一跳被取走。owner 已消失也不会永远飘着(`owner == ""` 由最低负载
节点兜底消费),与 Revoke 的处理一致。
执行侧的 owner 判定保留为第二道防线(双保险,两层各有测试覆盖)。
## 测试(又抓到一次假绿 + 一次死循环)
- TestRestartTaskReachesNonLocalOwner:非 owner 处理一 token 后任务必须仍在
token 里,随后 owner 处理时恰好应用 1 次。
★ 第一次写它时反复把**同一个 State 值**喂回 OnToken,导致死循环;改成按
真实环的走法(每跳喂一个新 token)后正常。
- TestRestartOnlyAppliedByOwner:把 peer 设成**最低负载节点**(否则泛用认领
分支根本不会触发,测了等于没测),断言它也不得应用。
★ 第一版 peer 不是最低负载节点 ⇒ 删掉 selection 分支后测试仍然绿,是假绿;
改为最低负载后,删分支 → TestRestartTaskReachesNonLocalOwner 变红。
- 其余:TestRestartTaskBypassesDuplicateClaimGuard、TestRestartFlagSurvives
TokenSerialization 保持绿。
★ 第三次「双向验证」的价值:一个测试抓不到,**另一个**抓到了。单靠一个测试
的绿就下结论是不安全的。
go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
@ -382,6 +382,24 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
|
||||
}
|
||||
continue // owned elsewhere — ride to the owner
|
||||
}
|
||||
if t.Restart {
|
||||
// Owner-directed, exactly like a revoke: only the node holding the
|
||||
// forward may act. Selecting it anywhere else would let a non-owner
|
||||
// spawn a worker for somebody else's forward (observed live: a
|
||||
// non-owner logged "restarted t4" and ran the forward itself).
|
||||
//
|
||||
// A restart whose forward is already gone is a no-op consumed by the
|
||||
// lowest node, so a task can never ride forever if its owner
|
||||
// departed before seeing it.
|
||||
if owner := e.state.TopologyOwner(t); owner == e.ID {
|
||||
target = t
|
||||
break
|
||||
} else if owner == "" && selfIsLowest {
|
||||
target = t // forward gone — consume and re-enable nothing
|
||||
break
|
||||
}
|
||||
continue // owned elsewhere — ride to the owner
|
||||
}
|
||||
if selfIsLowest {
|
||||
target = t // generic forward create — lowest-load claim
|
||||
break
|
||||
@ -505,20 +523,13 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
|
||||
}
|
||||
}
|
||||
if claimed.Restart {
|
||||
// Only the OWNER may act. Pending tasks are visible to every member, so
|
||||
// without this guard whichever node happened to process the task would
|
||||
// spawn a worker for a forward attributed to somebody else — an
|
||||
// orphaned worker plus a duplicate claim, exactly what the duplicate
|
||||
// guard above exists to prevent. (Observed live: a non-owner logged
|
||||
// "restarted t4" and ran the forward itself.)
|
||||
//
|
||||
// A restart whose owner has vanished is not an error: the entry is
|
||||
// re-enabled, OfflineReassign() will move it to pending on the next
|
||||
// departure sweep, and the normal claim path then re-homes it.
|
||||
// Only reached when this node owns the forward (or it is already gone
|
||||
// and we are the fallback consumer — see the selection loop).
|
||||
owner := e.state.TopologyOwner(claimed)
|
||||
e.state.UpdateTopologyDisabled(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort, false)
|
||||
if owner != e.ID {
|
||||
log.Printf("ring[%s] skip restart %s: owned by %s", e.ID, claimed.ID, owner)
|
||||
// Forward vanished before we got here: nothing to re-enable.
|
||||
log.Printf("ring[%s] discard restart %s: no owner", e.ID, claimed.ID)
|
||||
continue
|
||||
}
|
||||
if e.Handler != nil {
|
||||
|
||||
Reference in New Issue
Block a user