fix(cluster): 停用的转发会在重启后自己复活 + 令牌轮转日志刷屏

从线上三节点(192.168.2.{30,106,60})的日志里挖出四个问题,本轮修三个。

## 1. 停用的转发会复活(功能性缺陷,实测仍在发生)

线上现象:`minecraft` 在 store 里 disabled=1,worker 却仍在跑,今天
09:46 还在刷 `connect to local service [192.168.2.60:25565]: connection
refused` —— 对着一个用户刻意没启的本地服务死刷。同时三台 logs 里躺着
13956 / 3769 条 `proxy [x] already exists`,每 33 秒一轮。

四处叠加导致:
- stopForward 先读 link,再 SetLinkDisabled(true),然后把**改之前**的
  副本交给 RevokeTask ⇒ published task 带 disabled=false(实测 id/flag
  都对不上:topology 里 link.id=223,store 里同一行是 199)
- ClaimFn **无条件** SetLinkDisabled(...,false)。原意是"重新认领时清掉
  停用标记",但启动 reconcile 只要 worker 不在就重新 Claim ⇒ 每次重启
  都是"先清标记再起 worker"
- RevokeFn 停了 worker 却**没删 topology 条目**,条目活过 worker
- 于是下次重启 reconcile 看到"owned 但 worker 不在"→ 再次 Claim → 死循环

修法(把 disabled 的所有权交回两个用户动作):
- Claim **只读** disabled 决定要不要起 worker;为 true 时连 topology
  条目一起摘掉,绝不复活
- 启动 reconcile 先按 store 跳过 disabled 的条目(省掉无谓的
  claim→skip 往返)
- RevokeFn 除停 worker 外,同时 RemoveTopologyEntry —— 撤销必须是
  完整退役,不能只是"停一下"
- stopForward 把 Disabled=true 随 task 发布出去,让持有该转发的节点
  即使本地 store 行陈旧也能判断这次停用是用户主动的

## 2. 令牌轮转日志零信息量却占满磁盘

每轮固定 3 行(OnToken cycle=N / forward cycle=N / token-send -> 200),
2 轮/秒,实测本机 **355 行/分钟、7 天 357 万行**,把真事件全淹了。
同一份信息(cycle / lastSync / roundDelayMs / 成员存活)本来就能从
GET /api/manager/cluster/ring 结构化拿到。

加 W4F_DEBUG 开关(沿用项目既有 W4F_ 前缀约定):稳态三行降级为 debug、
默认关闭;**失败路径一律保留** —— 发送失败、陈旧令牌、非 2xx 正是别人
grep 的对象,静音它们是坏交易。实测同样 12 秒:36 行 → 4 行。

## 3. Link.ID 在 ReplaceLinks 之后必然失效

ReplaceLinks 是 DELETE + 重新 INSERT,sqlite 给每行**新的自增 id**。
任何在改写前捕获的 Link(典型:随 token 环跑的 Link)手里的 id 要么查无
此行,要么命中另一条转发 —— 实测捕获 alpha id=1,改写后新表是 3/4/5,
GetLink(1) 直接落空。

新增 LinkByTriple(local, remote, port) 按自然键查(业务代码本来就一律用
这个三元组标识转发),并把 claim/reconcile 切过去。查无行返回
(Link{}, false, nil) 而非 error:新建的转发没有行,应当照常启动。

## 4. homeagent_device 孤儿(已澄清,非独立缺陷)

它 disabled=1 且从不在 topology 里,是缺陷 1 的另一面(停用标记没进
token),随本次修复覆盖,无需单独处理。

## 测试

新增 4 个测试文件,重点是**验证测试本身抓得住 bug**:
- 临时回退 `ln.Disabled = true` 这行 → TestStopForwardPublishesDisabled-
  FlagInRevokeTask 如期变红,还原后变绿
- ⚠️ 第一版回归测试只断言 store 层,是**假绿**:newTestHandler 的 Ring
  为 nil,RevokeTask 那条(真正坏掉的)路根本没执行。补了带 ring 的
  newRingTestHandler,直接断言**发布出去的 task 上的 flag**
- LinkByTriple 在 ReplaceLinks 前后保持稳定;GetLink(id) 的失效被固化成
  一个可见的说明性测试
- 停用/start 往返、per-forward 停用不误伤兄弟转发
- 错误路径不静音、W4F_DEBUG 各种取值

go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
JianFeeeee
2026-09-26 10:02:55 +08:00
parent f4fb964734
commit 46e8bc3703
11 changed files with 757 additions and 8 deletions

View File

@ -175,13 +175,39 @@ func main() {
return err
}
// A re-claim after a forward-centric stop leaves the link flagged
// disabled (by RevokeFn); clear it so renderRemote renders the
// proxy back in. No-op for a fresh claim.
_ = st.SetLinkDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, false)
// disabled (by RevokeFn). Do NOT clear that flag here: the ring
// re-claims automatically (startup reconcile re-spawns any owned
// forward whose worker is missing), so clearing it on claim made
// "stopped" un-durable — every restart resurrected a forward the user
// had explicitly stopped, and it then failed forever against a local
// service that was intentionally not running.
//
// The flag is now owned by the two user-facing actions:
// startForward → SetLinkDisabled(false) before submitting the task
// stopForward → SetLinkDisabled(true) before revoking it
// Claim only READS it to decide whether to bring the worker up. A fresh
// claim of a link with no row still starts, because the lookup miss
// below is treated as "not disabled".
disabled := false
if existing, found, err := st.LinkByTriple(tk.Link.Local, tk.Link.Remote, tk.Link.RemotePort); err != nil {
return err
} else if found {
disabled = existing.Disabled
}
// Start (or restart) the per-forward worker for exactly this link.
// Each forward has its own frpc process (keyed by the forward
// triple); restarting only this key leaves sibling forwards'
// processes untouched.
if disabled {
// A stopped forward must also not linger in the topology: leaving
// the entry behind is what let the reconcile loop above keep
// re-claiming it on every restart.
if ring != nil {
ring.RemoveTopologyEntry(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
}
log.Printf("ring[%s] claim %s skipped: %s→%s:%d is disabled", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
return nil
}
if tk.Remote.Enabled {
key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
if _, has := pm.Status(key); has {
@ -205,6 +231,14 @@ func main() {
if _, running := pm.Status(key); running {
_ = pm.Stop(key)
}
// Drop the topology entry too. Leaving it behind meant the entry
// outlived the worker, and the next startup reconcile saw a
// "missing" worker for an owned forward and re-claimed it — which
// restarted a forward the user had explicitly stopped. Revoking
// must be a complete retirement, not just a stop.
if ring != nil {
ring.RemoveTopologyEntry(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
}
log.Printf("ring[%s] revoked task %s: %s→%s:%d", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
return nil
},
@ -437,6 +471,16 @@ func main() {
if t.OwnerID != selfID {
continue
}
// A forward the user stopped must stay stopped: skip it here so a
// restart does not re-spawn its worker. The ClaimFn enforces the
// same rule (belt and braces — this also avoids a pointless
// claim→skip round trip per disabled forward on every boot).
if ln, found, err := st.LinkByTriple(t.Local.Name, t.Remote.Name, t.Link.RemotePort); err != nil {
log.Printf("ring[%s] reconcile: lookup %s→%s:%d: %v", ring.ID, t.Local.Name, t.Remote.Name, t.Link.RemotePort, err)
continue
} else if found && ln.Disabled {
continue
}
key := process.WorkerKey(t.Local.Name, t.Remote.Name, t.Link.RemotePort)
if _, has := pm.Status(key); has {
continue // worker already running