mirror of
https://gitcode.com/JianFeeeee/webui4frpc.git
synced 2026-10-03 07:34:00 +00:00
fix(cluster): 停用的转发会在重启后自己复活 + 令牌轮转日志刷屏
从线上三节点(192.168.2.{30,106,60})的日志里挖出四个问题,本轮修三个。
## 1. 停用的转发会复活(功能性缺陷,实测仍在发生)
线上现象:`minecraft` 在 store 里 disabled=1,worker 却仍在跑,今天
09:46 还在刷 `connect to local service [192.168.2.60:25565]: connection
refused` —— 对着一个用户刻意没启的本地服务死刷。同时三台 logs 里躺着
13956 / 3769 条 `proxy [x] already exists`,每 33 秒一轮。
四处叠加导致:
- stopForward 先读 link,再 SetLinkDisabled(true),然后把**改之前**的
副本交给 RevokeTask ⇒ published task 带 disabled=false(实测 id/flag
都对不上:topology 里 link.id=223,store 里同一行是 199)
- ClaimFn **无条件** SetLinkDisabled(...,false)。原意是"重新认领时清掉
停用标记",但启动 reconcile 只要 worker 不在就重新 Claim ⇒ 每次重启
都是"先清标记再起 worker"
- RevokeFn 停了 worker 却**没删 topology 条目**,条目活过 worker
- 于是下次重启 reconcile 看到"owned 但 worker 不在"→ 再次 Claim → 死循环
修法(把 disabled 的所有权交回两个用户动作):
- Claim **只读** disabled 决定要不要起 worker;为 true 时连 topology
条目一起摘掉,绝不复活
- 启动 reconcile 先按 store 跳过 disabled 的条目(省掉无谓的
claim→skip 往返)
- RevokeFn 除停 worker 外,同时 RemoveTopologyEntry —— 撤销必须是
完整退役,不能只是"停一下"
- stopForward 把 Disabled=true 随 task 发布出去,让持有该转发的节点
即使本地 store 行陈旧也能判断这次停用是用户主动的
## 2. 令牌轮转日志零信息量却占满磁盘
每轮固定 3 行(OnToken cycle=N / forward cycle=N / token-send -> 200),
2 轮/秒,实测本机 **355 行/分钟、7 天 357 万行**,把真事件全淹了。
同一份信息(cycle / lastSync / roundDelayMs / 成员存活)本来就能从
GET /api/manager/cluster/ring 结构化拿到。
加 W4F_DEBUG 开关(沿用项目既有 W4F_ 前缀约定):稳态三行降级为 debug、
默认关闭;**失败路径一律保留** —— 发送失败、陈旧令牌、非 2xx 正是别人
grep 的对象,静音它们是坏交易。实测同样 12 秒:36 行 → 4 行。
## 3. Link.ID 在 ReplaceLinks 之后必然失效
ReplaceLinks 是 DELETE + 重新 INSERT,sqlite 给每行**新的自增 id**。
任何在改写前捕获的 Link(典型:随 token 环跑的 Link)手里的 id 要么查无
此行,要么命中另一条转发 —— 实测捕获 alpha id=1,改写后新表是 3/4/5,
GetLink(1) 直接落空。
新增 LinkByTriple(local, remote, port) 按自然键查(业务代码本来就一律用
这个三元组标识转发),并把 claim/reconcile 切过去。查无行返回
(Link{}, false, nil) 而非 error:新建的转发没有行,应当照常启动。
## 4. homeagent_device 孤儿(已澄清,非独立缺陷)
它 disabled=1 且从不在 topology 里,是缺陷 1 的另一面(停用标记没进
token),随本次修复覆盖,无需单独处理。
## 测试
新增 4 个测试文件,重点是**验证测试本身抓得住 bug**:
- 临时回退 `ln.Disabled = true` 这行 → TestStopForwardPublishesDisabled-
FlagInRevokeTask 如期变红,还原后变绿
- ⚠️ 第一版回归测试只断言 store 层,是**假绿**:newTestHandler 的 Ring
为 nil,RevokeTask 那条(真正坏掉的)路根本没执行。补了带 ring 的
newRingTestHandler,直接断言**发布出去的 task 上的 flag**
- LinkByTriple 在 ReplaceLinks 前后保持稳定;GetLink(id) 的失效被固化成
一个可见的说明性测试
- 停用/start 往返、per-forward 停用不误伤兄弟转发
- 错误路径不静音、W4F_DEBUG 各种取值
go build / go vet / go test ./... 全绿,gofmt 干净。
This commit is contained in:
@ -175,13 +175,39 @@ func main() {
|
||||
return err
|
||||
}
|
||||
// A re-claim after a forward-centric stop leaves the link flagged
|
||||
// disabled (by RevokeFn); clear it so renderRemote renders the
|
||||
// proxy back in. No-op for a fresh claim.
|
||||
_ = st.SetLinkDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, false)
|
||||
// disabled (by RevokeFn). Do NOT clear that flag here: the ring
|
||||
// re-claims automatically (startup reconcile re-spawns any owned
|
||||
// forward whose worker is missing), so clearing it on claim made
|
||||
// "stopped" un-durable — every restart resurrected a forward the user
|
||||
// had explicitly stopped, and it then failed forever against a local
|
||||
// service that was intentionally not running.
|
||||
//
|
||||
// The flag is now owned by the two user-facing actions:
|
||||
// startForward → SetLinkDisabled(false) before submitting the task
|
||||
// stopForward → SetLinkDisabled(true) before revoking it
|
||||
// Claim only READS it to decide whether to bring the worker up. A fresh
|
||||
// claim of a link with no row still starts, because the lookup miss
|
||||
// below is treated as "not disabled".
|
||||
disabled := false
|
||||
if existing, found, err := st.LinkByTriple(tk.Link.Local, tk.Link.Remote, tk.Link.RemotePort); err != nil {
|
||||
return err
|
||||
} else if found {
|
||||
disabled = existing.Disabled
|
||||
}
|
||||
// Start (or restart) the per-forward worker for exactly this link.
|
||||
// Each forward has its own frpc process (keyed by the forward
|
||||
// triple); restarting only this key leaves sibling forwards'
|
||||
// processes untouched.
|
||||
if disabled {
|
||||
// A stopped forward must also not linger in the topology: leaving
|
||||
// the entry behind is what let the reconcile loop above keep
|
||||
// re-claiming it on every restart.
|
||||
if ring != nil {
|
||||
ring.RemoveTopologyEntry(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
|
||||
}
|
||||
log.Printf("ring[%s] claim %s skipped: %s→%s:%d is disabled", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
|
||||
return nil
|
||||
}
|
||||
if tk.Remote.Enabled {
|
||||
key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
|
||||
if _, has := pm.Status(key); has {
|
||||
@ -205,6 +231,14 @@ func main() {
|
||||
if _, running := pm.Status(key); running {
|
||||
_ = pm.Stop(key)
|
||||
}
|
||||
// Drop the topology entry too. Leaving it behind meant the entry
|
||||
// outlived the worker, and the next startup reconcile saw a
|
||||
// "missing" worker for an owned forward and re-claimed it — which
|
||||
// restarted a forward the user had explicitly stopped. Revoking
|
||||
// must be a complete retirement, not just a stop.
|
||||
if ring != nil {
|
||||
ring.RemoveTopologyEntry(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
|
||||
}
|
||||
log.Printf("ring[%s] revoked task %s: %s→%s:%d", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
|
||||
return nil
|
||||
},
|
||||
@ -437,6 +471,16 @@ func main() {
|
||||
if t.OwnerID != selfID {
|
||||
continue
|
||||
}
|
||||
// A forward the user stopped must stay stopped: skip it here so a
|
||||
// restart does not re-spawn its worker. The ClaimFn enforces the
|
||||
// same rule (belt and braces — this also avoids a pointless
|
||||
// claim→skip round trip per disabled forward on every boot).
|
||||
if ln, found, err := st.LinkByTriple(t.Local.Name, t.Remote.Name, t.Link.RemotePort); err != nil {
|
||||
log.Printf("ring[%s] reconcile: lookup %s→%s:%d: %v", ring.ID, t.Local.Name, t.Remote.Name, t.Link.RemotePort, err)
|
||||
continue
|
||||
} else if found && ln.Disabled {
|
||||
continue
|
||||
}
|
||||
key := process.WorkerKey(t.Local.Name, t.Remote.Name, t.Link.RemotePort)
|
||||
if _, has := pm.Status(key); has {
|
||||
continue // worker already running
|
||||
|
||||
Reference in New Issue
Block a user