mirror of
https://gitcode.com/JianFeeeee/webui4frpc.git
synced 2026-10-03 23:53:59 +00:00
feat(cluster): 停用改为「标记」语义,让 disabled 真正随令牌环跨节点传播
承接用户提问「设计上停用不是本来就会跨节点传输吗」——核实结论:结构上确实 如此(TopoEntry.Link 是完整 store.Link,整个 State 随 token 每轮广播),但 实际路径断了。断点正是「撤销会删掉 topology 条目」:条目是 flag 的载体, 删了就无处传播,于是停用只能靠一次性 revoke 任务投递给 owner,**owner 当时 不在线就收不到**(实测 .60 记 disabled=1 / .106 记 0,就是这么来的)。 ## 改为标记而非移除 撤销不再 RemoveTopology,而是 UpdateTopologyDisabled(true),条目保留、 Link.Disabled=true、Active=false。Active 正是为此存在:OfflineReassign() 只处理 Active 条目,所以停用的转发在 owner 掉线时不会被重新排队。 - 新增 UpdateTopologyDisabled / TopologyDisabled(照 UpdateTopologyGroup 的桥) - 新增 store.ReconcileLinkDisabled 作接收端:adoption 时把环上的 flag 落进 本地 store;本节点没有该转发时补一条 disabled 占位行(否则日后在本节点被 claim 会复活),enable 则不建行 - SetTopologySync 由单向(store→环)扩为双向:群组仍上行,disabled 下行 - AddTopology 的 Active 跟随 Link.Disabled(原本硬编码 true,认领一个停用 转发就会复活它) - 审计日志细分 forward.stop / forward.start,与 forward.remove 区分 ## 语义变更带出的两个新问题(都已修) 1. **「启动」这条路断了**。条目保留 ⇒ SubmitTask 被去重挡下,而认领路径的 duplicate-claim 防御又会丢弃「已有 owner」的任务 ⇒ 重启任务发不出去,owner 永远收不到,转发**能停不能起**。 修:新增 Task.Restart 这一独立任务类型 + SubmitRestart + Handler.RestartFn, 显式绕过 duplicate-claim 防御并原地复活(不重复建条目、不重跑 claim 簿记)。 SubmitTask 的守卫同时从 HasTask 收窄为新的 HasActiveTask(跳过 disabled 条目 与撤销任务);saveCanvas 的判断相应改用 HasActiveTask,避免每次保存都对 已标记的转发重复发撤销。 2. 原本两处 RemoveTopologyEntry 调用(ClaimFn/RevokeFn 的 disabled 分支)在 新语义下会把本该保留的条目删掉,改为 UpdateTopologyDisabled。 ## 测试(每个都做了「回退修复行→必须变红→还原变绿」双向验证) - TestStoppedTopologyEntrySurvivesAdoption —— 离线成员也能学到停用, 一次性 revoke 任务永远做不到这一点 - TestStoppedForwardNotRequeuedOnNodeDeparture / TestAddTopologyRespectsDisabledFlag —— 标记而非删除为何安全 - TestSubmitTaskNotBlockedByStoppedEntry / TestSubmitTaskStillDedupesActiveForward - TestRestartTaskBypassesDuplicateClaimGuard / TestRestartFlagSurvivesTokenSerialization - TestStopThenStartPublishesRestartTask(HTTP 端到端,断言**任务真的发出**) - TestReconcileLinkDisabled*(store 侧三条) ★ 两次踩到**假绿**:第一版只断言 store 层(newTestHandler 的 Ring 为 nil, 坏掉的路根本没执行);第二版在 re-enable **之后**才调 SubmitTask,此时新旧 谓词结果相同,测不出差异。都是靠「回退修复行看是否变红」抓出来的 —— 这个 双向验证已经是本项目的固定动作。 go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
@ -18,6 +18,11 @@ import (
|
||||
type Handler interface {
|
||||
Claim(ctx context.Context, tk *Task) error
|
||||
Revoke(ctx context.Context, tk *Task) error
|
||||
// Restart re-enables an already-claimed forward and brings its worker back.
|
||||
// Needed because a stop keeps the topology entry, which makes both the
|
||||
// creation path (idempotency guard) and the duplicate-claim guard refuse to
|
||||
// re-create an existing forward.
|
||||
Restart(ctx context.Context, tk *Task) error
|
||||
RuntimeLoad() Load
|
||||
}
|
||||
|
||||
@ -390,23 +395,34 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
|
||||
break
|
||||
}
|
||||
if claimed.Revoke {
|
||||
// The revoke handler must run REGARDLESS of whether a topology entry
|
||||
// existed. RemoveTopology() returns false when the forward is not in
|
||||
// the topology, and gating the handler on it made the revoke a silent
|
||||
// no-op in exactly the case that matters: a forward the user stopped
|
||||
// while it was already absent from the topology kept its worker
|
||||
// running forever (observed live — a disabled forward kept dialing a
|
||||
// local service that was intentionally down, thousands of
|
||||
// connection-refused lines). Stopping the worker and retiring the
|
||||
// entry are independent duties; the entry may simply not be there.
|
||||
removed := e.state.RemoveTopology(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort)
|
||||
// A user-requested stop is recorded as a FLAG on the topology entry,
|
||||
// not as a removal. The entry is the carrier that makes the decision
|
||||
// travel: TopoEntry.Link is a full store.Link and State rides the
|
||||
// token every cycle, so keeping the entry is what lets a stop reach
|
||||
// every member — including a node that was offline when the stop was
|
||||
// issued. Removing the entry instead left the flag nowhere to live,
|
||||
// which is why a stop used to converge only if the owner happened to
|
||||
// be online to receive a one-shot revoke task.
|
||||
//
|
||||
// Marking Active=false is also what makes the entry inert for the
|
||||
// rebalancing paths: OfflineReassign() skips inactive entries, and
|
||||
// the startup reconcile skips them too, so a stopped forward is
|
||||
// neither re-queued on a node departure nor re-claimed on a restart.
|
||||
//
|
||||
// Handler.Revoke still runs unconditionally — it is what actually
|
||||
// stops the per-forward worker, and it must run even when there was no
|
||||
// entry to mark (gating it on RemoveTopology()'s result made stops of
|
||||
// already-absent forwards silently do nothing, so their workers ran
|
||||
// forever: observed live as thousands of connection-refused lines
|
||||
// against a local service that was intentionally down).
|
||||
marked := e.state.UpdateTopologyDisabled(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort, true)
|
||||
if e.Handler != nil {
|
||||
if err := e.Handler.Revoke(ctx, claimed); err != nil {
|
||||
log.Printf("ring[%s] revoke %s: %v", e.ID, claimed.ID, err)
|
||||
}
|
||||
}
|
||||
if removed && e.Log != nil {
|
||||
_, _ = e.Log.Append(e.ID, LogForwardRemove, map[string]any{
|
||||
if marked && e.Log != nil {
|
||||
_, _ = e.Log.Append(e.ID, LogForwardStop, map[string]any{
|
||||
"taskId": claimed.ID, "local": claimed.Local.Name, "remote": claimed.Remote.Name,
|
||||
})
|
||||
}
|
||||
@ -475,10 +491,38 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
|
||||
// different node. Safe for OfflineReassign: that path deletes the
|
||||
// topology entry BEFORE re-queueing, so TopologyOwner returns "" and
|
||||
// the legitimate re-claim passes through.
|
||||
if owner := e.state.TopologyOwner(claimed); owner != "" {
|
||||
log.Printf("ring[%s] drop duplicate task %s: %s→%s:%d already owned by %s",
|
||||
e.ID, claimed.ID, claimed.Local.Name, claimed.Remote.Name,
|
||||
claimed.Link.RemotePort, owner)
|
||||
//
|
||||
// A RESTART task is the deliberate exception this guard must not eat: it
|
||||
// exists precisely to re-activate a forward whose entry is still present
|
||||
// (a stop keeps the entry as the flag's carrier), so applying the guard
|
||||
// would swallow every re-enable and leave the forward stopped forever.
|
||||
if !claimed.Restart {
|
||||
if owner := e.state.TopologyOwner(claimed); owner != "" {
|
||||
log.Printf("ring[%s] drop duplicate task %s: %s→%s:%d already owned by %s",
|
||||
e.ID, claimed.ID, claimed.Local.Name, claimed.Remote.Name,
|
||||
claimed.Link.RemotePort, owner)
|
||||
continue
|
||||
}
|
||||
}
|
||||
if claimed.Restart {
|
||||
// Re-enable in place and hand the worker back to this node (the
|
||||
// entry's owner). This must NOT create a second topology entry, which
|
||||
// is why it skips the claim bookkeeping below.
|
||||
e.state.UpdateTopologyDisabled(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort, false)
|
||||
if e.Handler != nil {
|
||||
if err := e.Handler.Restart(ctx, claimed); err != nil {
|
||||
e.state.PendingTasks[claimed.ID] = claimed
|
||||
log.Printf("ring[%s] restart %s failed: %v", e.ID, claimed.ID, err)
|
||||
break
|
||||
}
|
||||
}
|
||||
if e.Log != nil {
|
||||
_, _ = e.Log.Append(e.ID, LogForwardStart, map[string]any{
|
||||
"taskId": claimed.ID, "local": claimed.Local.Name, "remote": claimed.Remote.Name,
|
||||
})
|
||||
}
|
||||
log.Printf("ring[%s] restarted %s: %s→%s:%d", e.ID, claimed.ID,
|
||||
claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort)
|
||||
continue
|
||||
}
|
||||
if e.Handler != nil {
|
||||
@ -817,14 +861,87 @@ func (e *Engine) RevokeTask(local store.Local, remote store.Remote, link store.L
|
||||
}
|
||||
|
||||
func (e *Engine) SubmitTask(local store.Local, remote store.Remote, link store.Link) *Task {
|
||||
if e.HasTask(local.Name, remote.Name, link.RemotePort) {
|
||||
// Guard on an ACTIVE task only. A topology entry that is merely marked
|
||||
// disabled is kept around as the carrier for the stop flag, so counting it
|
||||
// here would make this a permanent no-op and a stopped forward could never
|
||||
// be started again.
|
||||
if e.HasActiveTask(local.Name, remote.Name, link.RemotePort) {
|
||||
return nil
|
||||
}
|
||||
return e.state.AddPending(local, remote, link)
|
||||
}
|
||||
|
||||
// HasTask reports whether a forward with the same local/remote/remotePort is
|
||||
// already pending or active in the topology (idempotency guard for resaves).
|
||||
// HasTopologyEntry reports whether the forward has a topology entry (whether or
|
||||
// not it is marked disabled). A stop keeps the entry as the flag's carrier, so
|
||||
// "has an entry" is what distinguishes "already claimed by someone" from "never
|
||||
// submitted" — the distinction startForward needs.
|
||||
func (e *Engine) HasTopologyEntry(local, remote string, port int) bool {
|
||||
_, found := e.state.TopologyDisabled(local, remote, port)
|
||||
return found
|
||||
}
|
||||
|
||||
// SubmitRestart publishes a task asking the current owner of an already-claimed
|
||||
// forward to bring its worker back up.
|
||||
//
|
||||
// Neither of the existing channels can express "restart what already exists":
|
||||
// SubmitTask is guarded by HasActiveTask and the entry is still present, so it
|
||||
// dedupes; and the claim path drops any task for a forward that already has an
|
||||
// owner (its defence against duplicate-claim collisions). Re-enabling a stopped
|
||||
// forward therefore needs its own task kind, which the owner applies without
|
||||
// re-running the claim bookkeeping.
|
||||
func (e *Engine) SubmitRestart(local store.Local, remote store.Remote, link store.Link) *Task {
|
||||
// The task must not carry the stop: the whole point is to re-enable.
|
||||
link.Disabled = false
|
||||
t := &Task{
|
||||
ID: e.state.NextTaskID(),
|
||||
Local: local,
|
||||
Remote: remote,
|
||||
Link: link,
|
||||
Created: time.Now().Unix(),
|
||||
Restart: true,
|
||||
}
|
||||
if e.state.PendingTasks == nil {
|
||||
e.state.PendingTasks = map[string]*Task{}
|
||||
}
|
||||
e.state.PendingTasks[t.ID] = t
|
||||
return t
|
||||
}
|
||||
|
||||
// HasActiveTask reports whether a forward with this key is genuinely in flight
|
||||
// or running: a non-revocation pending task, or a topology entry that is not
|
||||
// marked disabled.
|
||||
//
|
||||
// It is deliberately narrower than HasTask. Under mark-don't-remove a stopped
|
||||
// forward keeps its topology entry (that entry is what carries the flag between
|
||||
// nodes), so "an entry exists" no longer means "this forward is active".
|
||||
// SubmitTask must use THIS predicate, otherwise starting a stopped forward —
|
||||
// startForward() publishes the enable and then calls SubmitTask — would be
|
||||
// swallowed by its own idempotency guard.
|
||||
func (e *Engine) HasActiveTask(local, remote string, port int) bool {
|
||||
for _, t := range e.state.PendingList() {
|
||||
if t.Revoke {
|
||||
continue // a revocation is the opposite of an active forward
|
||||
}
|
||||
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
|
||||
return true
|
||||
}
|
||||
}
|
||||
for _, t := range e.state.TopologyList() {
|
||||
if t.Link.Disabled {
|
||||
continue // stopped; the entry is only a flag carrier
|
||||
}
|
||||
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// HasTask reports whether ANY record for this forward exists — a pending task
|
||||
// (including an in-flight revocation) or a topology entry (including one marked
|
||||
// disabled). This is the "have we already handled this key" question, used to
|
||||
// avoid re-issuing work on repeated canvas saves. For "is it actually running"
|
||||
// use HasActiveTask instead.
|
||||
func (e *Engine) HasTask(local, remote string, port int) bool {
|
||||
for _, t := range e.state.PendingList() {
|
||||
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
|
||||
@ -846,6 +963,19 @@ func (e *Engine) UpdateTopologyGroup(local, remote string, port int, group strin
|
||||
return e.state.UpdateTopologyGroup(local, remote, port, group)
|
||||
}
|
||||
|
||||
// UpdateTopologyDisabled marks a forward enabled/disabled in the topology so
|
||||
// the decision rides the next token cycle to every member. Paired with
|
||||
// TopologyDisabled, which the adoption hook uses to learn peers' decisions.
|
||||
func (e *Engine) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
|
||||
return e.state.UpdateTopologyDisabled(local, remote, port, disabled)
|
||||
}
|
||||
|
||||
// TopologyDisabled reports the cluster's view of a forward's disabled flag.
|
||||
// The bool is false when the forward has no topology entry (nothing to learn).
|
||||
func (e *Engine) TopologyDisabled(local, remote string, port int) (bool, bool) {
|
||||
return e.state.TopologyDisabled(local, remote, port)
|
||||
}
|
||||
|
||||
// IsLeader reports whether this node is the current ring leader.
|
||||
func (e *Engine) IsLeader() bool { return e.state.LeaderID == e.ID }
|
||||
|
||||
|
||||
Reference in New Issue
Block a user