11 Commits
v0.1.1 ... main

Author SHA1 Message Date
918ca5d5ba fix(cluster): 停用状态以「环」为权威,消除长期离线导致的永久分歧
承接用户指出的遗留:links 表是每节点本地副本,靠「store→环上行 + 环→store
下行」双向 sync 收敛,但两边都可能被覆盖。其中「本地 store 权威」这条规则有
硬伤,本次改掉。

## 先复现,再动手

写探针验证「入环 token 会不会冲掉本地未传播的决定」,结果比预期严重:

    after adopting a stale token: known=true disabled=false
    LOCAL DISABLE WAS WIPED by an incoming token

OnToken/AdoptState 是 `e.state = tk.State` **整体替换**。两次 token 之间做出的
停用决定,只要下一轮到达的 token 是「决定之前」捕获的,就会被整个洗掉 ——
决定永远传不出去,转发照旧运行。Group 之所以看起来没这个问题,是因为它每次
adoption 都被 `SetTopologySync` 从 store 重新推上去,而 disabled 没有对应的
「尚未传播」保护。

## 改法:环权威 + 本地决定带「未确认」标记

1. **环权威**:adoption 时把环上的 disabled 写穿本地 store(write-through,
   不是监听器),任何 peer 的决定都在一个 token 周期内落地。这终结了旧规则
   「各人信自己那份」造成的永久分歧。顺序上只有单向要求:引擎先重新断言本地
   未确认决定,再跑 host sync,所以读环绝不会覆盖用户刚做的操作。

2. **未确认决定受保护**(RingEngine.localDisabled):`UpdateTopologyDisabled`
   记下决定,adoption 后由 `reconcileLocalDisabled` 重新断言到刚采纳的 state 上,
   于是它会随下一轮 token 传出去。环报回同值时删除该键(全cluster已一致);
   转发从 topology 消失时一并清扫,map 不会无限增长。**false 同样受保护** ——
   重新启用也需要传播,丢掉它会把转发永久留在停用态。

3. `TopologyDisabled()` 对未确认决定短路返回本地值,避免状态页与用户刚做的
   操作相反。

## 测试(又抓到一个假绿)

新增 4 条:停用/启用跨 adoption 存活、确认后停止断言(让位给 peer 的后续决定)、
追踪表不累积。

★ `TestLocalDisableSurvivesAdoption` **第一版是假绿**:它断言
`TopologyDisabled()`,而该访问器会短路到本地决定,于是「即便即将转发的 state
仍是 enabled」它也报 true。改成断言**下游节点会看到什么**(用一个 peer 引擎
AdoptState 本引擎的 state)后,去掉重新断言如期变红:

    the forwarded token carries disabled=false, want true

双向验证通过,这个教训要记:测「声明」而不是测「实际传播的状态」,等于没测。

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 11:44:28 +08:00
c8ab2584a8 chore(cluster): 去掉 restart 的重复日志行
实测 .30 在处理一个 restart 任务时打出两条完全相同的
`restarted t4: portal→aliyun-frps:4434`(同毫秒、同内容),一度像是执行了两次。
核实 worker 进程只有一个(restart 对已运行进程是幂等的),所以不是重复执行 ——
是引擎层(ring_engine.go)与应用层(RestartFn)各打了一条同样的日志。

保留应用层那条:它紧跟实际 worker 操作(谁真正把进程拉起来的),信息位置更准;
引擎层只保留 discard/no-owner 这类异常路径的日志,一次重启对应一行。

go build / go vet / go test ./... 全绿。
2026-09-26 11:06:14 +08:00
8ab52b238a fix(cluster): restart 任务必须定向投递给 owner,不能被非 owner 吞掉
上一提交(5cdc052)只加了执行侧 owner 判定,实测仍然失败:从 .60(非 owner)
启动 owner 在 .30 的 portal,.30 的 worker 一直没起来,且三台日志里既没有
`restarted` 也没有任何错误。

## 真因:任务被非 owner「消费」掉了

pending 命令由 runCommands 的 `for {}` 循环每轮重取 PendingList(),取出后
ClaimPending 即从 map 移除。我当时在「owner != e.ID」时把任务塞回
PendingTasks —— 于是它**立刻又变回待处理**,下一轮循环再次取到,无限
`defer restart`(单测直接跑成死循环,300s 超时)。

而上一版的 `continue` 同样是错的:ClaimPending 已经把任务移除,continue
等于消费,owner 永远收不到。

## 修法:放进「选择」循环,和 Revoke 完全同构

runCommands 里 Revoke 早就有正确的定向投递范式:
    owner == e.ID        → 我持有,执行
    owner == "" 且最低负载 → 转发已消失,兜底消费
    否则                  → continue(任务**留在 token 里**随环前进)
restart 照抄这套。非 owner 只是不选中它,任务随 token 传给下一个节点,直到
owner 那一跳被取走。owner 已消失也不会永远飘着(`owner == ""` 由最低负载
节点兜底消费),与 Revoke 的处理一致。

执行侧的 owner 判定保留为第二道防线(双保险,两层各有测试覆盖)。

## 测试(又抓到一次假绿 + 一次死循环)

- TestRestartTaskReachesNonLocalOwner:非 owner 处理一 token 后任务必须仍在
  token 里,随后 owner 处理时恰好应用 1 次。
  ★ 第一次写它时反复把**同一个 State 值**喂回 OnToken,导致死循环;改成按
  真实环的走法(每跳喂一个新 token)后正常。
- TestRestartOnlyAppliedByOwner:把 peer 设成**最低负载节点**(否则泛用认领
  分支根本不会触发,测了等于没测),断言它也不得应用。
  ★ 第一版 peer 不是最低负载节点 ⇒ 删掉 selection 分支后测试仍然绿,是假绿;
  改为最低负载后,删分支 → TestRestartTaskReachesNonLocalOwner 变红。
- 其余:TestRestartTaskBypassesDuplicateClaimGuard、TestRestartFlagSurvives
  TokenSerialization 保持绿。

★ 第三次「双向验证」的价值:一个测试抓不到,**另一个**抓到了。单靠一个测试
的绿就下结论是不安全的。

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 11:03:00 +08:00
5cdc052d01 fix(cluster): restart 任务必须只由 owner 执行
上线实测抓到的:从 .60(非 owner)启动 owner 在 .30 的 portal,日志出现
**两条** `.60 restarted t4` —— 关键帧是 `[192.168.2.60:7500] restarted t4`。
pending 任务对所有成员可见,而 restart 分支没有 owner 判定,于是谁先轮到
谁就执行:非 owner 给自己的机器起了一个属于别人的转发 worker,而真正的
owner(.30)什么也没做,worker 一直没起来。

这正是上面 duplicate-claim 防御本来要防的「孤儿 worker + 重复认领」,只是
restart 分支为了绕开那层防御,把 owner 校验也一并跳过了 —— 绕开的是
「重复建条目」的必要性,不是「只有 owner 能动手」的必要性。

修法与 RemoveNode 分支一致:`owner != e.ID` 时只清标志、不执行 handler。
owner 已消失的情况不是错误:条目已被重新标为启用,OfflineReassign() 会在
下一次离线清理时把它转入 pending,再由正常认领流程重新安置。

测试 TestRestartOnlyAppliedByOwner:同一份 restart 任务分别交给 owner 与非
owner,断言非 owner 调用 0 次、owner 恰好 1 次。双向验证 —— 去掉守卫后
如期变红("a NON-owner applied the restart 1 time(s)"),还原后变绿。

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 10:48:28 +08:00
1c835425de feat(cluster): 停用改为「标记」语义,让 disabled 真正随令牌环跨节点传播
承接用户提问「设计上停用不是本来就会跨节点传输吗」——核实结论:结构上确实
如此(TopoEntry.Link 是完整 store.Link,整个 State 随 token 每轮广播),但
实际路径断了。断点正是「撤销会删掉 topology 条目」:条目是 flag 的载体,
删了就无处传播,于是停用只能靠一次性 revoke 任务投递给 owner,**owner 当时
不在线就收不到**(实测 .60 记 disabled=1 / .106 记 0,就是这么来的)。

## 改为标记而非移除

撤销不再 RemoveTopology,而是 UpdateTopologyDisabled(true),条目保留、
Link.Disabled=true、Active=false。Active 正是为此存在:OfflineReassign()
只处理 Active 条目,所以停用的转发在 owner 掉线时不会被重新排队。

- 新增 UpdateTopologyDisabled / TopologyDisabled(照 UpdateTopologyGroup 的桥)
- 新增 store.ReconcileLinkDisabled 作接收端:adoption 时把环上的 flag 落进
  本地 store;本节点没有该转发时补一条 disabled 占位行(否则日后在本节点被
  claim 会复活),enable 则不建行
- SetTopologySync 由单向(store→环)扩为双向:群组仍上行,disabled 下行
- AddTopology 的 Active 跟随 Link.Disabled(原本硬编码 true,认领一个停用
  转发就会复活它)
- 审计日志细分 forward.stop / forward.start,与 forward.remove 区分

## 语义变更带出的两个新问题(都已修)

1. **「启动」这条路断了**。条目保留 ⇒ SubmitTask 被去重挡下,而认领路径的
   duplicate-claim 防御又会丢弃「已有 owner」的任务 ⇒ 重启任务发不出去,owner
   永远收不到,转发**能停不能起**。
   修:新增 Task.Restart 这一独立任务类型 + SubmitRestart + Handler.RestartFn,
   显式绕过 duplicate-claim 防御并原地复活(不重复建条目、不重跑 claim 簿记)。
   SubmitTask 的守卫同时从 HasTask 收窄为新的 HasActiveTask(跳过 disabled 条目
   与撤销任务);saveCanvas 的判断相应改用 HasActiveTask,避免每次保存都对
   已标记的转发重复发撤销。

2. 原本两处 RemoveTopologyEntry 调用(ClaimFn/RevokeFn 的 disabled 分支)在
   新语义下会把本该保留的条目删掉,改为 UpdateTopologyDisabled。

## 测试(每个都做了「回退修复行→必须变红→还原变绿」双向验证)

- TestStoppedTopologyEntrySurvivesAdoption —— 离线成员也能学到停用,
  一次性 revoke 任务永远做不到这一点
- TestStoppedForwardNotRequeuedOnNodeDeparture / TestAddTopologyRespectsDisabledFlag
  —— 标记而非删除为何安全
- TestSubmitTaskNotBlockedByStoppedEntry / TestSubmitTaskStillDedupesActiveForward
- TestRestartTaskBypassesDuplicateClaimGuard / TestRestartFlagSurvivesTokenSerialization
- TestStopThenStartPublishesRestartTask(HTTP 端到端,断言**任务真的发出**)
- TestReconcileLinkDisabled*(store 侧三条)

★ 两次踩到**假绿**:第一版只断言 store 层(newTestHandler 的 Ring 为 nil,
坏掉的路根本没执行);第二版在 re-enable **之后**才调 SubmitTask,此时新旧
谓词结果相同,测不出差异。都是靠「回退修复行看是否变红」抓出来的 —— 这个
双向验证已经是本项目的固定动作。

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 10:44:04 +08:00
041cc04dd6 fix(cluster): revoke 在无 topology 条目时静默失效 + 停用状态跨节点不同步
上一提交(46e8bc3)上线后实测:`.106` 重启后,被用户停用的
`minecraft` worker **仍然被拉起**。继续挖出两处,均为静默失效型。

## 1. Handler.Revoke 被 RemoveTopology 的返回值挡住了

    if claimed.Revoke {
        if e.state.RemoveTopology(...) {      // ← false 时整个块跳过
            if e.Handler != nil { e.Handler.Revoke(...) }
        }
    }

停 worker 的唯一动作被关在「topology 里确实有条目」的条件里。而
RemoveTopology 对**不在 topology 的转发返回 false** —— 也就是说,越是
需要清理的僵尸转发(条目已丢、worker 还在),revoke 越是完全不做。

这正是 minecraft 的处境:停用请求到达时它已不在 topology ⇒ 返回 false
⇒ Revoke 不执行 ⇒ worker 永远不停 ⇒ 每 33 秒刷一次 connection refused。

修法:停 worker 与摘条目是**两件独立的事**,条目不在也照样停。
(审计日志 forward.remove 仍只在真的摘掉条目时写,这是对的。)

## 2. 停用状态是节点本地的,从不跨节点同步

`links` 表每节点各一份。实测同一时刻 `.60` 记 disabled=1、`.106` 记
disabled=0 —— 因为只有收到停用请求的那个节点写了 store,而**持有 worker
的 owner 往往是另一台机器**,它的副本仍是"启用"。Claim 只读本地 store
⇒ owner 认为该转发是启用的 ⇒ 又被拉起。

修法:task 已携带 Disabled,以它为准,RevokeFn 在本节点把它落库
(已有行改写;没有行则补一条 disabled 记录,防止日后在本节点被 claim
时复活)。

## 测试

- TestRevokeIdempotent 原先只断言"不 panic",正好漏掉这个 bug —— 它
  允许「Handler.Revoke 从不调用」通过。现补上断言:拓扑里没有条目时
  **仍必须调用** Handler.Revoke。
- 验证该测试有效:把 `&& removed` gate 加回去 → 如期变红;还原后变绿。

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 10:05:55 +08:00
46e8bc3703 fix(cluster): 停用的转发会在重启后自己复活 + 令牌轮转日志刷屏
从线上三节点(192.168.2.{30,106,60})的日志里挖出四个问题,本轮修三个。

## 1. 停用的转发会复活(功能性缺陷,实测仍在发生)

线上现象:`minecraft` 在 store 里 disabled=1,worker 却仍在跑,今天
09:46 还在刷 `connect to local service [192.168.2.60:25565]: connection
refused` —— 对着一个用户刻意没启的本地服务死刷。同时三台 logs 里躺着
13956 / 3769 条 `proxy [x] already exists`,每 33 秒一轮。

四处叠加导致:
- stopForward 先读 link,再 SetLinkDisabled(true),然后把**改之前**的
  副本交给 RevokeTask ⇒ published task 带 disabled=false(实测 id/flag
  都对不上:topology 里 link.id=223,store 里同一行是 199)
- ClaimFn **无条件** SetLinkDisabled(...,false)。原意是"重新认领时清掉
  停用标记",但启动 reconcile 只要 worker 不在就重新 Claim ⇒ 每次重启
  都是"先清标记再起 worker"
- RevokeFn 停了 worker 却**没删 topology 条目**,条目活过 worker
- 于是下次重启 reconcile 看到"owned 但 worker 不在"→ 再次 Claim → 死循环

修法(把 disabled 的所有权交回两个用户动作):
- Claim **只读** disabled 决定要不要起 worker;为 true 时连 topology
  条目一起摘掉,绝不复活
- 启动 reconcile 先按 store 跳过 disabled 的条目(省掉无谓的
  claim→skip 往返)
- RevokeFn 除停 worker 外,同时 RemoveTopologyEntry —— 撤销必须是
  完整退役,不能只是"停一下"
- stopForward 把 Disabled=true 随 task 发布出去,让持有该转发的节点
  即使本地 store 行陈旧也能判断这次停用是用户主动的

## 2. 令牌轮转日志零信息量却占满磁盘

每轮固定 3 行(OnToken cycle=N / forward cycle=N / token-send -> 200),
2 轮/秒,实测本机 **355 行/分钟、7 天 357 万行**,把真事件全淹了。
同一份信息(cycle / lastSync / roundDelayMs / 成员存活)本来就能从
GET /api/manager/cluster/ring 结构化拿到。

加 W4F_DEBUG 开关(沿用项目既有 W4F_ 前缀约定):稳态三行降级为 debug、
默认关闭;**失败路径一律保留** —— 发送失败、陈旧令牌、非 2xx 正是别人
grep 的对象,静音它们是坏交易。实测同样 12 秒:36 行 → 4 行。

## 3. Link.ID 在 ReplaceLinks 之后必然失效

ReplaceLinks 是 DELETE + 重新 INSERT,sqlite 给每行**新的自增 id**。
任何在改写前捕获的 Link(典型:随 token 环跑的 Link)手里的 id 要么查无
此行,要么命中另一条转发 —— 实测捕获 alpha id=1,改写后新表是 3/4/5,
GetLink(1) 直接落空。

新增 LinkByTriple(local, remote, port) 按自然键查(业务代码本来就一律用
这个三元组标识转发),并把 claim/reconcile 切过去。查无行返回
(Link{}, false, nil) 而非 error:新建的转发没有行,应当照常启动。

## 4. homeagent_device 孤儿(已澄清,非独立缺陷)

它 disabled=1 且从不在 topology 里,是缺陷 1 的另一面(停用标记没进
token),随本次修复覆盖,无需单独处理。

## 测试

新增 4 个测试文件,重点是**验证测试本身抓得住 bug**:
- 临时回退 `ln.Disabled = true` 这行 → TestStopForwardPublishesDisabled-
  FlagInRevokeTask 如期变红,还原后变绿
- ⚠️ 第一版回归测试只断言 store 层,是**假绿**:newTestHandler 的 Ring
  为 nil,RevokeTask 那条(真正坏掉的)路根本没执行。补了带 ring 的
  newRingTestHandler,直接断言**发布出去的 task 上的 flag**
- LinkByTriple 在 ReplaceLinks 前后保持稳定;GetLink(id) 的失效被固化成
  一个可见的说明性测试
- 停用/start 往返、per-forward 停用不误伤兄弟转发
- 错误路径不静音、W4F_DEBUG 各种取值

go build / go vet / go test ./... 全绿,gofmt 干净。
2026-09-26 10:02:55 +08:00
f4fb964734 fix(web): 窄屏布局三处根因 + 页面滚动塌陷
在 375/320 屏上实机取证(三节点集群 192.168.2.{30,106,60}),修掉四个
独立缺陷。前三个是布局根因,第四个是被前三个掩盖的真 bug。

1. 全站没有 box-sizing 重置
   content-box 下 `.sidebar { width:100%; padding:10px 12px }` 在 375 视口
   实算 399px(320 屏 344px)。顶栏主题色块右缘 332 > 320,被
   `.sidebar{overflow-x:hidden}` 静默裁掉 ⇒ 主题切换在手机上点不到。
   桌面侧栏同样是 274px 而非声明的 246px。
   注:已有两个「窄屏顶栏右侧溢出」commit(64d3359 用 flex min-width/vw
   限宽、0559cc5 改 backdrop-filter)都在治症状,真因是盒模型。

2. fit-view-on-init 把画布缩到不可读
   画布用绝对图形坐标(x=60/760,两列随转发数向下增长),bbox 远大于
   手机视口。实测合成缩放 375 屏 0.3135 ⇒ 248px 的卡片渲成 85px、
   有效字号 4.39px;**1280 桌面也只有 0.6011 / 8.42px**(5 条转发即
   触发),所以这不只是窄屏问题。
   改为删掉 fit-view-on-init 自己接管:MIN_ZOOM=0.65 下限 +
   frameCanvas() 把最左上节点重贴到内边距边缘(fitView 是居中的,窄屏
   下会把两列接缝摆在视口正中,两边各看一半)。duration:0 保证量 rect
   时动画已停;初始 frame 只认一次,避免后续 add 节点把正在平移的用户
   拽走。挂 @init + @nodes-initialized(store 与节点分先后到达,任一
   单独触发都会 fit 空画布)。

3. el-dialog 固定宽度超出手机视口
   Element Plus 把 width 写成内联 --el-dialog-width。375 屏实测
   left=0 right=420,右上关闭按钮在屏外,只能横向滚 overlay 才够得到。
   窄屏下 clamp 到 calc(100vw - 24px)。

4. 【滚动塌陷】状态页在手机上完全无法滑动
   `.status-page` 是 height:100% + overflow:hidden 的 flex 列,内层
   `.status-body` 靠 flex:1 + overflow-y:auto 自己滚。375 屏下可用高度
   651px,而两个不可压缩子元素 topbar(122) + kpis(565) 已占 687px ⇒
   `.status-body` 被分到 **0px**,其中却有 1682px 内容。同时外层
   `.content` scrollHeight==clientHeight(无可滚内容)⇒ 整页没有任何
   可滚动元素,手势完全无响应。
   修法:让 `.content`(App.vue 已声明 flex:1 + overflow-y:auto)做唯一
   滚动容器,页面自身不再 height:100%/overflow。ClusterView /
   SettingsView 同类嵌套一并去掉,从根上消灭「按剩余空间算高度」这类
   算术错误。另把 ≤380px 的 KPI 单列改回两列(单列四张全宽卡自身就
   565px,超过整个视口,是塌陷的起因)。

验证(构建产物挂反代打真后端 + 共享 Chromium CDP):
- 缩放 0.31→0.65、卡片 85→177px、有效字号 4.39→9.1px(320/375/414/1280 一致)
- bodyScrollW 344→320(320 屏),横向溢出清零;主题色块全部落回屏内
- 对话框 left 12 / right 363(完全在屏内,关闭键可达)
- 9 种视口 × 5 个页面 = 45 格全部「末块可达 + 无横向滚动」,console 零报错
  修复前:状态页 375/320/414 全部无滚动;1280 桌面状态/设置同样中招
- 三节点滚动重启部署,集群收敛 pending=0 / topo=4,owner 与实际
  worker 进程一一对应,journal 无 panic/fatal,reclaimed=2 各节点自愈
2026-09-18 11:00:35 +08:00
64d3359981 fix(web): 窄屏顶栏右侧溢出 — flex 自动最小尺寸 + vw 限宽两处误算
两个原因叠加导致顶栏横向溢出:

1. "webui4frpc" 是不可断行单词,.sb-brand-text 没给 min-width:0 +
   overflow:hidden,其 flex 自动最小尺寸就等于 min-content(≈85px),
   顶栏因此拿不出空间给右侧身份区,超出部分溢出。现给品牌文字整条链路
   加可收缩约束,标题用 ellipsis 截断。

2. .sb-identity 用 max-width:42vw / .id-name 用 14vw 限宽:vw 含垂直
   滚动条宽度,比实际可用内容宽度大;且 max-width 只封顶盒子,内部
   flex:0 0 auto 的角色徽标与退出按钮仍会把它撑开。改为全链路
   min-width:0 + overflow:hidden,由 flex 自行分配(.sb-foot 也从
   flex:0 0 auto 改 0 1 auto,允许被压缩)。

另加 .sidebar { overflow-x: hidden } 作兜底,任何子项算错都不再产生
横向滚动条;≤380px 隐藏角色徽标(它不可收缩,"超级管理员"≈70px)。

验证:tsc --noEmit 通过、npm run build 通过、产物 CSS 已确认规则生效;
三节点部署 0.1.2 后环正常(cycle 推进、pending=0、topo=5)。
2026-09-03 08:50:23 +08:00
0559cc55bd fix(web): 窄屏底部导航栏错位到顶部 — 顶栏 backdrop-filter 创建了包含块
.sb-nav 用 position:fixed;bottom:0 想贴视口底部,但父元素 .sidebar 带
backdrop-filter。backdrop-filter 与 transform/filter 同理,会为后代的
position:fixed 创建【包含块】—— 于是 bottom:0 变成相对顶栏下沿定位而不是
视口,tab bar 贴在顶栏正下方,看起来就是固定在屏幕顶部。

窄屏下关掉顶栏的 backdrop-filter,改用近实色底
color-mix(var(--w4f-card-solid) 92%, transparent);玻璃模糊保留在
.sb-nav 自身(元素自己的 backdrop-filter 不影响自己的定位)。

顺带 bump 到 0.1.2(version 包 / release.sh / build_installers.sh /
rpm spec + changelog / README 安装示例文件名)。

验证:tsc --noEmit 通过、npm run build 通过;产物 CSS 中确认 720px 段的
.sidebar 已含 backdrop-filter:none;三节点部署后 /status 报告 0.1.2,
环正常(cycle 推进、pending=0、5 条转发 frpc 进程与 topology 归属一致)。
2026-09-03 08:45:22 +08:00
609f0df517 chore: upload_assets.py 默认包含原生安装包
默认文件筛选原来只匹配 tar.gz/zip/SHA256SUMS,deb/rpm/pkg/setup.exe
必须逐个显式传参才会上传(v0.1.0 与 v0.1.1 都是分两趟传的)。
改为按发布产物后缀白名单筛选;用 -setup.exe 而非裸 .exe,避免误收
bin/ 下的裸二进制。
2026-09-03 07:30:53 +08:00
30 changed files with 2263 additions and 88 deletions

View File

@ -79,7 +79,7 @@ cd web && npm run build && rm -rf internal/httpapi/dist && cp -r web/dist intern
### Linux (Debian / Ubuntu) ### Linux (Debian / Ubuntu)
```bash ```bash
sudo dpkg -i webui4frpc-0.1.1-linux-amd64.deb # 或 linux-arm64.deb sudo dpkg -i webui4frpc-0.1.2-linux-amd64.deb # 或 linux-arm64.deb
sudo nano /etc/webui4frpc/webui4frpc.env # 改掉默认密码 sudo nano /etc/webui4frpc/webui4frpc.env # 改掉默认密码
sudo systemctl restart webui4frpc sudo systemctl restart webui4frpc
``` ```
@ -87,7 +87,7 @@ sudo systemctl restart webui4frpc
### Linux (RHEL / Fedora / openEuler / 默认 / 龙蜥) ### Linux (RHEL / Fedora / openEuler / 默认 / 龙蜥)
```bash ```bash
sudo rpm -i webui4frpc-0.1.1-linux-x86_64.rpm sudo rpm -i webui4frpc-0.1.2-linux-x86_64.rpm
sudo vi /etc/webui4frpc/webui4frpc.env sudo vi /etc/webui4frpc/webui4frpc.env
sudo systemctl enable --now webui4frpc sudo systemctl enable --now webui4frpc
``` ```
@ -97,7 +97,7 @@ sudo systemctl enable --now webui4frpc
### Windows ### Windows
双击 `webui4frpc-0.1.1-windows-amd64-setup.exe`(ARM 设备用 `windows-arm64-setup.exe`)。 双击 `webui4frpc-0.1.2-windows-amd64-setup.exe`(ARM 设备用 `windows-arm64-setup.exe`)。
安装向导会询问监听地址与管理员凭据,完成后: 安装向导会询问监听地址与管理员凭据,完成后:
- 程序装到 `C:\Program Files\webui4frpc` - 程序装到 `C:\Program Files\webui4frpc`
@ -107,8 +107,8 @@ sudo systemctl enable --now webui4frpc
### macOS ### macOS
```bash ```bash
sudo installer -pkg webui4frpc-0.1.1-darwin-arm64.pkg -target / # Apple Silicon sudo installer -pkg webui4frpc-0.1.2-darwin-arm64.pkg -target / # Apple Silicon
sudo installer -pkg webui4frpc-0.1.1-darwin-amd64.pkg -target / # Intel sudo installer -pkg webui4frpc-0.1.2-darwin-amd64.pkg -target / # Intel
``` ```
或直接双击 `.pkg` 走图形安装器。安装后 launchd 开机自启,配置在 或直接双击 `.pkg` 走图形安装器。安装后 launchd 开机自启,配置在
@ -126,7 +126,7 @@ sudo launchctl kickstart -k system/com.jianfeeeee.webui4frpc
不想装服务只想跑一下,用对应的 `.tar.gz` / `.zip`,解压即用: 不想装服务只想跑一下,用对应的 `.tar.gz` / `.zip`,解压即用:
```bash ```bash
tar xzf webui4frpc-0.1.1-linux-amd64.tar.gz tar xzf webui4frpc-0.1.2-linux-amd64.tar.gz
./webui4frpc -addr 127.0.0.1:7500 -user admin -password 你的密码 -workdir ./data ./webui4frpc -addr 127.0.0.1:7500 -user admin -password 你的密码 -workdir ./data
``` ```

View File

@ -175,13 +175,41 @@ func main() {
return err return err
} }
// A re-claim after a forward-centric stop leaves the link flagged // A re-claim after a forward-centric stop leaves the link flagged
// disabled (by RevokeFn); clear it so renderRemote renders the // disabled (by RevokeFn). Do NOT clear that flag here: the ring
// proxy back in. No-op for a fresh claim. // re-claims automatically (startup reconcile re-spawns any owned
_ = st.SetLinkDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, false) // forward whose worker is missing), so clearing it on claim made
// "stopped" un-durable — every restart resurrected a forward the user
// had explicitly stopped, and it then failed forever against a local
// service that was intentionally not running.
//
// The flag is now owned by the two user-facing actions:
// startForward → SetLinkDisabled(false) before submitting the task
// stopForward → SetLinkDisabled(true) before revoking it
// Claim only READS it to decide whether to bring the worker up. A fresh
// claim of a link with no row still starts, because the lookup miss
// below is treated as "not disabled".
disabled := false
if existing, found, err := st.LinkByTriple(tk.Link.Local, tk.Link.Remote, tk.Link.RemotePort); err != nil {
return err
} else if found {
disabled = existing.Disabled
}
// Start (or restart) the per-forward worker for exactly this link. // Start (or restart) the per-forward worker for exactly this link.
// Each forward has its own frpc process (keyed by the forward // Each forward has its own frpc process (keyed by the forward
// triple); restarting only this key leaves sibling forwards' // triple); restarting only this key leaves sibling forwards'
// processes untouched. // processes untouched.
if disabled {
// Mark the entry stopped rather than deleting it. The entry is the
// carrier that keeps the flag travelling around the ring, and
// UpdateTopologyDisabled also clears Active — which is what makes
// this entry inert for OfflineReassign() and the startup
// reconcile, so it is not resurrected later.
if ring != nil {
ring.UpdateTopologyDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, true)
}
log.Printf("ring[%s] claim %s skipped: %s→%s:%d is disabled", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
return nil
}
if tk.Remote.Enabled { if tk.Remote.Enabled {
key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort) key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
if _, has := pm.Status(key); has { if _, has := pm.Status(key); has {
@ -201,13 +229,103 @@ func main() {
// stop/start: it wiped sibling forwards sharing the local or remote. // stop/start: it wiped sibling forwards sharing the local or remote.
RevokeFn: func(ctx context.Context, tk *cluster.Task) error { RevokeFn: func(ctx context.Context, tk *cluster.Task) error {
_ = st.SetLinkDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, true) _ = st.SetLinkDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, true)
// The links table is node-local: only the node that received the
// stop request has the flag persisted, while the node that OWNS the
// forward is the one holding the worker (and is often a different
// machine). Trusting only the local row meant the owner's copy could
// still read disabled=false. The task now carries the flag, so use it
// as the authority and write it here.
ln, found, err := st.LinkByTriple(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
switch {
case err != nil:
return err
case found:
ln.Disabled = true
// Persist via the same rewrite path the canvas uses, so the flag
// survives a restart on THIS node too.
links, err := st.ListLinks()
if err != nil {
return err
}
for i := range links {
if links[i].Local == ln.Local && links[i].Remote == ln.Remote && links[i].RemotePort == ln.RemotePort {
links[i].Disabled = true
}
}
if err := st.ReplaceLinks(links); err != nil {
return err
}
case tk.Link.Disabled:
// No local row yet (this node never claimed the forward) but the
// task asserts the stop — record it so a later claim on this node
// cannot resurrect the forward.
loc, rem := tk.Local, tk.Remote
links, err := st.ListLinks()
if err != nil {
return err
}
links = append(links, store.Link{
Local: loc.Name, Remote: rem.Name, RemotePort: tk.Link.RemotePort,
OffsetX: tk.Link.OffsetX, OffsetY: tk.Link.OffsetY,
Group: tk.Link.Group, Disabled: true,
})
if err := st.ReplaceLinks(links); err != nil {
return err
}
}
key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort) key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
if _, running := pm.Status(key); running { if _, running := pm.Status(key); running {
_ = pm.Stop(key) _ = pm.Stop(key)
} }
// Mark the entry stopped instead of deleting it: the entry carries
// the flag around the ring, and clearing Active keeps it inert for
// the rebalancing paths (OfflineReassign skips inactive entries, and
// the startup reconcile skips disabled forwards), so the worker is
// not re-spawned on the next restart. Deleting it here is what used
// to leave the stop with no carrier at all.
if ring != nil {
ring.UpdateTopologyDisabled(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort, true)
}
log.Printf("ring[%s] revoked task %s: %s→%s:%d", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort) log.Printf("ring[%s] revoked task %s: %s→%s:%d", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
return nil return nil
}, },
// RestartFn brings an already-claimed forward back after a stop. It
// must NOT run the claim bookkeeping (ClaimFn), which appends links and
// writes topology state: the forward is already established, and a stop
// deliberately keeps its topology entry, so all that is needed is to
// clear the flag and respawn the worker. The engine has already
// re-marked the topology entry enabled before calling this.
RestartFn: func(ctx context.Context, tk *cluster.Task) error {
// Clear the stop on this node's own store, using the same path the
// start handler uses, so the local copy cannot disagree with the
// ring and block a later restart.
if ln, found, err := st.LinkByTriple(tk.Link.Local, tk.Link.Remote, tk.Link.RemotePort); err != nil {
return err
} else if found && ln.Disabled {
links, err := st.ListLinks()
if err != nil {
return err
}
for i := range links {
if links[i].Local == ln.Local && links[i].Remote == ln.Remote && links[i].RemotePort == ln.RemotePort {
links[i].Disabled = false
}
}
if err := st.ReplaceLinks(links); err != nil {
return err
}
}
if tk.Remote.Enabled {
key := process.WorkerKey(tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
if _, has := pm.Status(key); has {
_ = pm.Restart(key)
} else {
_ = pm.Start(key)
}
}
log.Printf("ring[%s] restarted %s: %s→%s:%d", selfID, tk.ID, tk.Local.Name, tk.Remote.Name, tk.Link.RemotePort)
return nil
},
}, },
ringTransport.SendTo(func(nodeID string) string { ringTransport.SendTo(func(nodeID string) string {
if ring == nil { if ring == nil {
@ -232,10 +350,24 @@ func main() {
// left does NOT auto-rejoin. // left does NOT auto-rejoin.
ring.SetPeerPersist(func(peersJSON string) error { return st.SetClusterPeers(peersJSON) }) ring.SetPeerPersist(func(peersJSON string) error { return st.SetClusterPeers(peersJSON) })
// Re-apply local store group labels onto the ring topology after each // Re-apply local store state onto the ring topology after each state
// state adoption. Without this, group changes made via HTTP handlers // adoption, and learn the ring's disabled decisions back into the store.
//
// Without the group half, group changes made via HTTP handlers
// (POST /forwards/assign) are overwritten by the next e.state = tk.State // (POST /forwards/assign) are overwritten by the next e.state = tk.State
// and never propagate to other nodes. // and never propagate.
//
// The disabled half makes the RING authoritative for stop/start. That is the
// only source all members can agree on: the links table is a per-node copy, so
// treating it as authoritative let two nodes disagree indefinitely (observed
// live: one node reading disabled=1 while the owner still read 0 and kept its
// worker running). This is a write-through, not a listener — it runs on every
// adoption, so any peer's decision lands here within a token cycle.
//
// Ordering matters in one direction only: the engine has already re-applied
// this node's not-yet-confirmed local decisions onto the adopted state (see
// reconcileLocalDisabled), so reading the ring here can never clobber an
// action the user just took on this node.
ring.SetTopologySync(func() { ring.SetTopologySync(func() {
links, err := st.ListLinks() links, err := st.ListLinks()
if err != nil { if err != nil {
@ -244,6 +376,16 @@ func main() {
for _, ln := range links { for _, ln := range links {
ring.UpdateTopologyGroup(ln.Local, ln.Remote, ln.RemotePort, ln.Group) ring.UpdateTopologyGroup(ln.Local, ln.Remote, ln.RemotePort, ln.Group)
} }
for _, te := range ring.State().TopologyList() {
disabled, known := ring.TopologyDisabled(te.Local.Name, te.Remote.Name, te.Link.RemotePort)
if !known {
continue
}
if ln, found, err := st.LinkByTriple(te.Local.Name, te.Remote.Name, te.Link.RemotePort); err == nil && found && ln.Disabled == disabled {
continue // store already agrees with the ring
}
_ = st.ReconcileLinkDisabled(te.Local.Name, te.Remote.Name, te.Link.RemotePort, disabled)
}
}) })
h := &httpapi.Handler{ h := &httpapi.Handler{
@ -437,6 +579,16 @@ func main() {
if t.OwnerID != selfID { if t.OwnerID != selfID {
continue continue
} }
// A forward the user stopped must stay stopped: skip it here so a
// restart does not re-spawn its worker. The ClaimFn enforces the
// same rule (belt and braces — this also avoids a pointless
// claim→skip round trip per disabled forward on every boot).
if ln, found, err := st.LinkByTriple(t.Local.Name, t.Remote.Name, t.Link.RemotePort); err != nil {
log.Printf("ring[%s] reconcile: lookup %s→%s:%d: %v", ring.ID, t.Local.Name, t.Remote.Name, t.Link.RemotePort, err)
continue
} else if found && ln.Disabled {
continue
}
key := process.WorkerKey(t.Local.Name, t.Remote.Name, t.Link.RemotePort) key := process.WorkerKey(t.Local.Name, t.Remote.Name, t.Link.RemotePort)
if _, has := pm.Status(key); has { if _, has := pm.Status(key); has {
continue // worker already running continue // worker already running

91
internal/cluster/debug.go Normal file
View File

@ -0,0 +1,91 @@
package cluster
import (
"log"
"os"
"sync/atomic"
)
// logf is the package's logging seam: every cluster log line funnels through it
// so the debug gate in debug.go has a single place to hook.
func logf(format string, args ...any) { log.Printf(format, args...) }
// Token rotation is the ring's heartbeat: with a 2s round delay it fires
// continuously, and the default three log lines per round
// ("OnToken cycle=N" / "forward cycle=N to X" / "token-send ... -> 200")
// carry no information — no cycle number, address, timing or payload ever
// changes in the steady state. Measured on the live 3-node cluster that was
// ~510k lines/day on one node (3.5M lines in a week), which drowns every real
// event in the journal and fills the disk for a signal that is already
// available in structured form via GET /api/manager/cluster/ring (cycle,
// lastSync, roundDelayMs, node aliveness).
//
// So the steady-state lines are demoted to a debug level, off by default and
// enabled with W4F_DEBUG=token (or 1/all/true for every debug line). Failure
// paths are NOT demoted: a send error, a stale token, a timeout or a leader
// change is exactly what someone is grepping for, and losing those to a quiet
// default would be a bad trade.
const debugEnv = "W4F_DEBUG"
// debugToken logs a per-token-round heartbeat line. Suppressed unless
// W4F_DEBUG selects "token" (or a catch-all value).
func debugToken(format string, args ...any) {
if debugTokenOn.Load() {
logf(format, args...)
}
}
// debugAll logs an ad-hoc diagnostic line. Suppressed unless W4F_DEBUG is set
// to a catch-all value (1/all/true/*).
func debugAll(format string, args ...any) {
if debugAllOn.Load() {
logf(format, args...)
}
}
var (
debugTokenOn atomic.Bool
debugAllOn atomic.Bool
)
func init() { ReloadDebug() }
// ReloadDebug re-reads W4F_DEBUG. Called once at init so tests can flip it
// without restarting, and available at runtime for an operator who wants to
// watch the ring without a redeploy.
func ReloadDebug() {
v := os.Getenv(debugEnv)
switch normalized := normalizeDebugValue(v); normalized {
case "token":
debugTokenOn.Store(true)
debugAllOn.Store(false)
case "all":
debugTokenOn.Store(true)
debugAllOn.Store(true)
default:
debugTokenOn.Store(false)
debugAllOn.Store(false)
}
}
func normalizeDebugValue(v string) string {
// Compare case-insensitively without pulling in strings just for this.
out := make([]rune, 0, len(v))
for _, r := range v {
if r >= 'A' && r <= 'Z' {
r += 'a' - 'A'
}
out = append(out, r)
}
s := string(out)
switch s {
case "":
return ""
case "token", "tokens", "ring":
return "token"
}
// Any other non-empty value is a deliberate request for more output, so it
// is treated as a catch-all rather than silently muting the operator who
// set it. "0" lands here too: it was asked for, so honour it.
return "all"
}

View File

@ -0,0 +1,108 @@
package cluster
import (
"bytes"
"log"
"os"
"strings"
"testing"
)
// captureLog redirects the standard logger into a buffer for the duration of
// fn and returns what was written.
func captureLog(t *testing.T, fn func()) string {
t.Helper()
var buf bytes.Buffer
orig := log.Writer()
origFlags := log.Flags()
log.SetOutput(&buf)
log.SetFlags(0)
defer func() {
log.SetOutput(orig)
log.SetFlags(origFlags)
}()
fn()
return buf.String()
}
func TestDebugTokenSuppressedByDefault(t *testing.T) {
os.Unsetenv(debugEnv)
ReloadDebug()
out := captureLog(t, func() {
debugToken("ring[%s] OnToken cycle=%d", "node:7500", 42)
debugToken("ring[%s] forward cycle=%d to %s", "node:7500", 42, "next:7500")
debugToken("token-send %s: size=%dB elapsed=%v -> %d", "next:7500", 5605, "170ms", 200)
})
if out != "" {
t.Fatalf("steady-state token lines logged with W4F_DEBUG unset: %q", out)
}
}
func TestDebugTokenEnabledByEnv(t *testing.T) {
for _, v := range []string{"token", "TOKEN", "tokens", "ring", "1", "true", "yes", "all", "*", "yes-please"} {
t.Run(v, func(t *testing.T) {
t.Setenv(debugEnv, v)
ReloadDebug()
if !debugTokenOn.Load() {
t.Fatalf("W4F_DEBUG=%q should enable the token heartbeat", v)
}
out := captureLog(t, func() {
debugToken("ring[%s] OnToken cycle=%d", "node:7500", 7)
})
if !strings.Contains(out, "OnToken cycle=7") {
t.Fatalf("W4F_DEBUG=%q: expected the line to be logged, got %q", v, out)
}
})
}
}
// The whole point of the gate is to shrink the journal, so the failure paths
// must stay loud without any env var — losing a send error to a quiet default
// would be a bad trade.
func TestFailurePathsStayLoud(t *testing.T) {
os.Unsetenv(debugEnv)
ReloadDebug()
out := captureLog(t, func() {
logf("token-send %s: size=%dB elapsed=%v err=%v", "down:7500", 10, "5ms", "connection refused")
})
if !strings.Contains(out, "connection refused") {
t.Fatalf("send errors must never be gated: got %q", out)
}
}
func TestDebugAllOffForTokenOnly(t *testing.T) {
t.Setenv(debugEnv, "token")
ReloadDebug()
if !debugTokenOn.Load() {
t.Fatal("token level should enable debugToken")
}
if debugAllOn.Load() {
t.Fatal("token level must not enable the catch-all debugAll")
}
out := captureLog(t, func() { debugAll("scratch diagnostic") })
if out != "" {
t.Fatalf("debugAll should stay off at token level, got %q", out)
}
}
func TestNormalizeDebugValue(t *testing.T) {
cases := map[string]string{
"": "",
"token": "token",
"TOKEN": "token",
"Ring": "token",
"1": "all",
"true": "all",
"ALL": "all",
"*": "all",
"anything": "all", // unrecognised but deliberate → don't silence it
"0": "all", // 0 is still a deliberate request for output
}
for in, want := range cases {
if got := normalizeDebugValue(in); got != want {
t.Errorf("normalizeDebugValue(%q) = %q, want %q", in, got, want)
}
}
}

View File

@ -62,6 +62,12 @@ type Task struct {
// owning node cancels it (stop worker, drop from topology). Reuses the // owning node cancels it (stop worker, drop from topology). Reuses the
// same publish channel as creation (round-1 inject, round-2 apply). // same publish channel as creation (round-1 inject, round-2 apply).
Revoke bool `json:"revoke,omitempty"` Revoke bool `json:"revoke,omitempty"`
// Restart marks a RE-ENABLE task: the forward is already claimed and its
// topology entry still exists (a stop keeps the entry as the flag's
// carrier), so the creation path would dedupe the submission and the
// duplicate-claim guard would discard it. The owner applies this one
// unconditionally: re-mark the entry enabled and (re)spawn the worker.
Restart bool `json:"restart,omitempty"`
// RemoveNode: node ID to remove from the ring; the node self-removes when // RemoveNode: node ID to remove from the ring; the node self-removes when
// the command reaches it via the token. // the command reaches it via the token.
RemoveNode string `json:"removeNode,omitempty"` RemoveNode string `json:"removeNode,omitempty"`
@ -333,10 +339,15 @@ func (s *State) AddTopology(t *Task, ownerID string) *TopoEntry {
if s.Topology == nil { if s.Topology == nil {
s.Topology = map[string]*TopoEntry{} s.Topology = map[string]*TopoEntry{}
} }
// Active mirrors the task's disabled flag rather than being unconditionally
// true. A claim can legitimately carry a stopped forward (e.g. a node
// re-claiming from a departed peer), and forcing Active=true there would
// resurrect it: OfflineReassign only re-queues Active entries, and every
// reader that filters on Active would treat it as running again.
e := &TopoEntry{ e := &TopoEntry{
TaskID: t.ID, OwnerID: ownerID, TaskID: t.ID, OwnerID: ownerID,
Local: t.Local, Remote: t.Remote, Link: t.Link, Local: t.Local, Remote: t.Remote, Link: t.Link,
Active: true, Active: !t.Link.Disabled,
} }
s.Topology[t.ID] = e s.Topology[t.ID] = e
return e return e
@ -408,6 +419,39 @@ func (s *State) ForwardsOwnedBy(owner string) []*TopoEntry {
// (local, remote, port) triple. Returns true if found. The updated entry // (local, remote, port) triple. Returns true if found. The updated entry
// propagates to all nodes via the next token cycle — group changes sync // propagates to all nodes via the next token cycle — group changes sync
// through the ring without a dedicated command. // through the ring without a dedicated command.
// UpdateTopologyDisabled sets the disabled flag on the topology entry matching
// the forward's natural key. Mirror of UpdateTopologyGroup: it makes a locally
// decided stop/start part of the topology so the next token cycle carries it to
// every other member (the alternative — a one-shot revoke task addressed at the
// owner — only converges if that owner happens to be online right then).
//
// Returns false when the forward is not in the topology, which is not an error:
// a stopped forward may legitimately have no entry yet.
func (s *State) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {
e.Link.Disabled = disabled
// Keep Active consistent with the flag so readers that key off it
// (status page, load accounting) agree with Link.Disabled.
e.Active = !disabled
return true
}
}
return false
}
// TopologyDisabled returns the disabled flag the cluster currently holds for a
// forward, and whether such an entry exists. Used by the adoption path to learn
// a peer's decision into the local store.
func (s *State) TopologyDisabled(local, remote string, port int) (bool, bool) {
for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {
return e.Link.Disabled, true
}
}
return false, false
}
func (s *State) UpdateTopologyGroup(local, remote string, port int, group string) bool { func (s *State) UpdateTopologyGroup(local, remote string, port int, group string) bool {
for _, e := range s.Topology { for _, e := range s.Topology {
if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port { if e.Local.Name == local && e.Remote.Name == remote && e.Link.RemotePort == port {

View File

@ -0,0 +1,157 @@
package cluster
import (
"context"
"testing"
"webui4frpc/internal/store"
)
// staleTokenWith copies the engine's current state but rewrites the given
// forward's disabled flag, standing in for a token a neighbour captured before
// this node's decision.
func staleTokenWith(eng *Engine, local, remote string, port int, disabled bool) State {
s := eng.state
out := State{
LeaderID: s.LeaderID,
Nodes: s.Nodes,
Cycle: s.Cycle,
Topology: map[string]*TopoEntry{},
}
for id, e := range s.Topology {
cp := *e
if cp.Local.Name == local && cp.Remote.Name == remote && cp.Link.RemotePort == port {
cp.Link.Disabled = disabled
cp.Active = !disabled
}
out.Topology[id] = &cp
}
return out
}
// TestLocalDisableSurvivesAdoption is the regression test for the wipe that made
// the ring's authority unachievable.
//
// OnToken/AdoptState replace e.state wholesale. A stop decided between two token
// cycles was therefore erased by the next adoption whenever the incoming token
// had been captured before the decision — so the flag never reached the other
// members and the forward kept running. Written as a probe BEFORE the fix: it
// reported "after adopting a stale token: known=true disabled=false".
func TestLocalDisableSurvivesAdoption(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state, SentAt: 1}); err != nil {
t.Fatal(err)
}
// Local decision: stop it.
eng.UpdateTopologyDisabled("web", "frps1", 18081, true)
if d, _ := eng.TopologyDisabled("web", "frps1", 18081); !d {
t.Fatal("precondition: the local stop should be visible")
}
// A token captured BEFORE the decision arrives.
stale := staleTokenWith(eng, "web", "frps1", 18081, false)
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: stale, SentAt: 2}); err != nil {
t.Fatal(err)
}
// Assert on the STATE, not the TopologyDisabled() accessor: the accessor
// short-circuits to the pending local decision, so it would report the stop
// even when the state about to be forwarded still says "enabled". What
// matters is that the token this node sends on carries the decision.
assertForwardedDisabled(t, eng, "web", "frps1", 18081, true)
}
// assertForwardedDisabled checks the disabled value a DOWNSTREAM node would see
// in the token this engine forwards — the value that actually propagates.
func assertForwardedDisabled(t *testing.T, eng *Engine, local, remote string, port int, want bool) {
t.Helper()
// Model one hop: a fresh engine adopts exactly what this one would send.
peer := newTestEngine("peer", false)
peer.AdoptState(eng.state)
got, known := peer.TopologyDisabled(local, remote, port)
if !known {
t.Fatalf("peer learned no entry for %s→%s:%d", local, remote, port)
}
if got != want {
t.Fatalf("the forwarded token carries disabled=%v, want %v — the decision did not "+
"propagate, so the other members would keep the old state", got, want)
}
}
// TestLocalReEnableSurvivesAdoption: a re-enable needs the same protection. The
// value false is easy to overlook — nothing "looks" stopped, yet losing it just
// as silently strands the forward in the stopped state.
func TestLocalReEnableSurvivesAdoption(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state, SentAt: 1}); err != nil {
t.Fatal(err)
}
eng.UpdateTopologyDisabled("web", "frps1", 18081, true)
eng.UpdateTopologyDisabled("web", "frps1", 18081, false)
// A token that still says disabled (captured mid-stop) arrives.
mid := staleTokenWith(eng, "web", "frps1", 18081, true)
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: mid, SentAt: 2}); err != nil {
t.Fatal(err)
}
assertForwardedDisabled(t, eng, "web", "frps1", 18081, false)
}
// TestConfirmedDecisionStopsBeingAsserted: once the ring reports the same value
// the decision is confirmed everywhere, so this node must stop overriding —
// otherwise it would fight a later decision made elsewhere.
func TestConfirmedDecisionStopsBeingAsserted(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state, SentAt: 1}); err != nil {
t.Fatal(err)
}
eng.UpdateTopologyDisabled("web", "frps1", 18081, true)
if len(eng.localDisabled) != 1 {
t.Fatalf("a pending decision should be tracked, got %v", eng.localDisabled)
}
// The ring comes back agreeing.
agreed := staleTokenWith(eng, "web", "frps1", 18081, true)
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: agreed, SentAt: 2}); err != nil {
t.Fatal(err)
}
if len(eng.localDisabled) != 0 {
t.Fatalf("a confirmed decision must stop being asserted, still tracking %v", eng.localDisabled)
}
// A peer now decides the opposite; this node must follow the ring.
peerSaysEnabled := staleTokenWith(eng, "web", "frps1", 18081, false)
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 3, State: peerSaysEnabled, SentAt: 3}); err != nil {
t.Fatal(err)
}
if d, _ := eng.TopologyDisabled("web", "frps1", 18081); d {
t.Fatal("this node kept overriding a peer's later decision; the ring is not authoritative")
}
}
// TestLocalDecisionsDoNotAccumulate: the tracking map is swept against the live
// topology, so a long-lived cluster cannot grow it without bound.
func TestLocalDecisionsDoNotAccumulate(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state, SentAt: 1}); err != nil {
t.Fatal(err)
}
// Decisions for forwards this node has never seen.
eng.UpdateTopologyDisabled("ghost1", "frps1", 1, true)
eng.UpdateTopologyDisabled("ghost2", "frps1", 2, true)
if len(eng.localDisabled) != 2 {
t.Fatalf("expected 2 tracked decisions, got %d", len(eng.localDisabled))
}
// A cycle runs: neither ghost is in the topology, so both are swept.
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state, SentAt: 2}); err != nil {
t.Fatal(err)
}
if len(eng.localDisabled) != 0 {
t.Fatalf("decisions for forwards outside the topology must be swept, got %v", eng.localDisabled)
}
}

View File

@ -7,6 +7,7 @@ import (
"encoding/json" "encoding/json"
"fmt" "fmt"
"log" "log"
"strconv"
"sync" "sync"
"time" "time"
@ -18,6 +19,11 @@ import (
type Handler interface { type Handler interface {
Claim(ctx context.Context, tk *Task) error Claim(ctx context.Context, tk *Task) error
Revoke(ctx context.Context, tk *Task) error Revoke(ctx context.Context, tk *Task) error
// Restart re-enables an already-claimed forward and brings its worker back.
// Needed because a stop keeps the topology entry, which makes both the
// creation path (idempotency guard) and the duplicate-claim guard refuse to
// re-create an existing forward.
Restart(ctx context.Context, tk *Task) error
RuntimeLoad() Load RuntimeLoad() Load
} }
@ -50,6 +56,24 @@ type Engine struct {
// group changes made via HTTP handlers between token cycles are overwritten // group changes made via HTTP handlers between token cycles are overwritten
// by the next state adoption and never propagate to other nodes. // by the next state adoption and never propagate to other nodes.
topologySync func() topologySync func()
// localDisabled records this node's own enabled/disabled decisions that have
// not yet been confirmed by the ring. It is what makes the RING authoritative
// without losing a decision that was made locally and has not yet had a
// chance to reach the other members.
//
// The problem it solves: OnToken/AdoptState do `e.state = tk.State`, a
// wholesale replacement. A stop decided between two token cycles therefore
// disappears on the very next adoption if that token was captured before the
// decision — the flag never reaches the rest of the ring, and the forward
// keeps running. (Verified with a probe before writing this: adopting a
// stale token flipped disabled back to false.)
//
// A key stays in this map until an adoption reports it disabled, at which
// point the whole cluster agrees and the key is dropped. Keys for forwards
// that vanish from the topology are swept on every update, so the map cannot
// grow without bound. The value false is what keeps a RE-ENABLE alive for the
// same reason a stop needs protecting.
localDisabled map[string]bool
state State state State
// myAddr maps our Node ID to the address peers dial. // myAddr maps our Node ID to the address peers dial.
@ -182,7 +206,9 @@ func (e *Engine) OnToken(ctx context.Context, tk *Token) (*Token, error) {
if tk.SentAt > e.lastTokenAt { if tk.SentAt > e.lastTokenAt {
e.lastTokenAt = tk.SentAt e.lastTokenAt = tk.SentAt
} }
log.Printf("ring[%s] OnToken cycle=%d", e.ID, tk.Cycle) // Per-round heartbeat — steady state, no information. See debug.go: this
// fired ~2x/second and dominated the journal on every node.
debugToken("ring[%s] OnToken cycle=%d", e.ID, tk.Cycle)
// Parallel rhythm timer: operations run while the pace clock ticks. // Parallel rhythm timer: operations run while the pace clock ticks.
// Delay scales with alive node count (more nodes → lower per-hop delay, // Delay scales with alive node count (more nodes → lower per-hop delay,
@ -229,6 +255,11 @@ func (e *Engine) OnToken(ctx context.Context, tk *Token) (*Token, error) {
// Re-apply local store overrides (group labels, disabled flags) onto // Re-apply local store overrides (group labels, disabled flags) onto
// the freshly adopted topology so they survive state adoption and // the freshly adopted topology so they survive state adoption and
// propagate to all nodes via the next token forward. // propagate to all nodes via the next token forward.
// Re-assert decisions this node made locally that the ring has not yet
// confirmed, BEFORE the host's topology sync runs: the sync reads
// TopologyDisabled to learn the cluster's view, so the local decision must
// already be applied or a stop could be reported as "enabled" and dropped.
e.reconcileLocalDisabled()
if e.topologySync != nil { if e.topologySync != nil {
e.topologySync() e.topologySync()
} }
@ -375,6 +406,24 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
} }
continue // owned elsewhere — ride to the owner continue // owned elsewhere — ride to the owner
} }
if t.Restart {
// Owner-directed, exactly like a revoke: only the node holding the
// forward may act. Selecting it anywhere else would let a non-owner
// spawn a worker for somebody else's forward (observed live: a
// non-owner logged "restarted t4" and ran the forward itself).
//
// A restart whose forward is already gone is a no-op consumed by the
// lowest node, so a task can never ride forever if its owner
// departed before seeing it.
if owner := e.state.TopologyOwner(t); owner == e.ID {
target = t
break
} else if owner == "" && selfIsLowest {
target = t // forward gone — consume and re-enable nothing
break
}
continue // owned elsewhere — ride to the owner
}
if selfIsLowest { if selfIsLowest {
target = t // generic forward create — lowest-load claim target = t // generic forward create — lowest-load claim
break break
@ -388,18 +437,37 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
break break
} }
if claimed.Revoke { if claimed.Revoke {
if e.state.RemoveTopology(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort) { // A user-requested stop is recorded as a FLAG on the topology entry,
// not as a removal. The entry is the carrier that makes the decision
// travel: TopoEntry.Link is a full store.Link and State rides the
// token every cycle, so keeping the entry is what lets a stop reach
// every member — including a node that was offline when the stop was
// issued. Removing the entry instead left the flag nowhere to live,
// which is why a stop used to converge only if the owner happened to
// be online to receive a one-shot revoke task.
//
// Marking Active=false is also what makes the entry inert for the
// rebalancing paths: OfflineReassign() skips inactive entries, and
// the startup reconcile skips them too, so a stopped forward is
// neither re-queued on a node departure nor re-claimed on a restart.
//
// Handler.Revoke still runs unconditionally — it is what actually
// stops the per-forward worker, and it must run even when there was no
// entry to mark (gating it on RemoveTopology()'s result made stops of
// already-absent forwards silently do nothing, so their workers ran
// forever: observed live as thousands of connection-refused lines
// against a local service that was intentionally down).
marked := e.state.UpdateTopologyDisabled(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort, true)
if e.Handler != nil { if e.Handler != nil {
if err := e.Handler.Revoke(ctx, claimed); err != nil { if err := e.Handler.Revoke(ctx, claimed); err != nil {
log.Printf("ring[%s] revoke %s: %v", e.ID, claimed.ID, err) log.Printf("ring[%s] revoke %s: %v", e.ID, claimed.ID, err)
} }
} }
if e.Log != nil { if marked && e.Log != nil {
_, _ = e.Log.Append(e.ID, LogForwardRemove, map[string]any{ _, _ = e.Log.Append(e.ID, LogForwardStop, map[string]any{
"taskId": claimed.ID, "local": claimed.Local.Name, "remote": claimed.Remote.Name, "taskId": claimed.ID, "local": claimed.Local.Name, "remote": claimed.Remote.Name,
}) })
} }
}
continue continue
} }
if claimed.RemoveNode != "" && claimed.RemoveNode == e.ID { if claimed.RemoveNode != "" && claimed.RemoveNode == e.ID {
@ -465,12 +533,46 @@ func (e *Engine) runCommands(ctx context.Context, tk *Token) error {
// different node. Safe for OfflineReassign: that path deletes the // different node. Safe for OfflineReassign: that path deletes the
// topology entry BEFORE re-queueing, so TopologyOwner returns "" and // topology entry BEFORE re-queueing, so TopologyOwner returns "" and
// the legitimate re-claim passes through. // the legitimate re-claim passes through.
//
// A RESTART task is the deliberate exception this guard must not eat: it
// exists precisely to re-activate a forward whose entry is still present
// (a stop keeps the entry as the flag's carrier), so applying the guard
// would swallow every re-enable and leave the forward stopped forever.
if !claimed.Restart {
if owner := e.state.TopologyOwner(claimed); owner != "" { if owner := e.state.TopologyOwner(claimed); owner != "" {
log.Printf("ring[%s] drop duplicate task %s: %s→%s:%d already owned by %s", log.Printf("ring[%s] drop duplicate task %s: %s→%s:%d already owned by %s",
e.ID, claimed.ID, claimed.Local.Name, claimed.Remote.Name, e.ID, claimed.ID, claimed.Local.Name, claimed.Remote.Name,
claimed.Link.RemotePort, owner) claimed.Link.RemotePort, owner)
continue continue
} }
}
if claimed.Restart {
// Only reached when this node owns the forward (or it is already gone
// and we are the fallback consumer — see the selection loop).
owner := e.state.TopologyOwner(claimed)
e.state.UpdateTopologyDisabled(claimed.Local.Name, claimed.Remote.Name, claimed.Link.RemotePort, false)
if owner != e.ID {
// Forward vanished before we got here: nothing to re-enable.
log.Printf("ring[%s] discard restart %s: no owner", e.ID, claimed.ID)
continue
}
if e.Handler != nil {
if err := e.Handler.Restart(ctx, claimed); err != nil {
e.state.PendingTasks[claimed.ID] = claimed
log.Printf("ring[%s] restart %s failed: %v", e.ID, claimed.ID, err)
break
}
}
if e.Log != nil {
_, _ = e.Log.Append(e.ID, LogForwardStart, map[string]any{
"taskId": claimed.ID, "local": claimed.Local.Name, "remote": claimed.Remote.Name,
})
}
// Logging is the app's job (RestartFn reports the actual worker
// outcome); the ring layer stays quiet so one restart is one line
// rather than two identical ones.
continue
}
if e.Handler != nil { if e.Handler != nil {
if err := e.Handler.Claim(ctx, claimed); err != nil { if err := e.Handler.Claim(ctx, claimed); err != nil {
e.state.PendingTasks[claimed.ID] = claimed e.state.PendingTasks[claimed.ID] = claimed
@ -724,6 +826,11 @@ func (e *Engine) AdoptState(s State) {
for id, t := range kept { for id, t := range kept {
e.state.PendingTasks[id] = t e.state.PendingTasks[id] = t
} }
// Re-assert decisions this node made locally that the ring has not yet
// confirmed, BEFORE the host's topology sync runs: the sync reads
// TopologyDisabled to learn the cluster's view, so the local decision must
// already be applied or a stop could be reported as "enabled" and dropped.
e.reconcileLocalDisabled()
if e.topologySync != nil { if e.topologySync != nil {
e.topologySync() e.topologySync()
} }
@ -792,6 +899,14 @@ func (e *Engine) RemoveNode(nodeID string) *Task {
return e.state.AddRemoveNode(nodeID) return e.state.AddRemoveNode(nodeID)
} }
// RemoveTopologyEntry drops the active topology entry for a forward identified
// by its natural key. Exposed so the app's claim path can retire an entry it
// refuses to serve (see ClaimFn: a disabled forward must stop looking active,
// otherwise the startup reconcile keeps re-claiming it on every restart).
func (e *Engine) RemoveTopologyEntry(local, remote string, port int) bool {
return e.state.RemoveTopology(local, remote, port)
}
// RevokeTask publishes a revocation for an established forward through the // RevokeTask publishes a revocation for an established forward through the
// same token channel; the owning node stops the worker and drops topology. // same token channel; the owning node stops the worker and drops topology.
func (e *Engine) RevokeTask(local store.Local, remote store.Remote, link store.Link) *Task { func (e *Engine) RevokeTask(local store.Local, remote store.Remote, link store.Link) *Task {
@ -799,14 +914,87 @@ func (e *Engine) RevokeTask(local store.Local, remote store.Remote, link store.L
} }
func (e *Engine) SubmitTask(local store.Local, remote store.Remote, link store.Link) *Task { func (e *Engine) SubmitTask(local store.Local, remote store.Remote, link store.Link) *Task {
if e.HasTask(local.Name, remote.Name, link.RemotePort) { // Guard on an ACTIVE task only. A topology entry that is merely marked
// disabled is kept around as the carrier for the stop flag, so counting it
// here would make this a permanent no-op and a stopped forward could never
// be started again.
if e.HasActiveTask(local.Name, remote.Name, link.RemotePort) {
return nil return nil
} }
return e.state.AddPending(local, remote, link) return e.state.AddPending(local, remote, link)
} }
// HasTask reports whether a forward with the same local/remote/remotePort is // HasTopologyEntry reports whether the forward has a topology entry (whether or
// already pending or active in the topology (idempotency guard for resaves). // not it is marked disabled). A stop keeps the entry as the flag's carrier, so
// "has an entry" is what distinguishes "already claimed by someone" from "never
// submitted" — the distinction startForward needs.
func (e *Engine) HasTopologyEntry(local, remote string, port int) bool {
_, found := e.state.TopologyDisabled(local, remote, port)
return found
}
// SubmitRestart publishes a task asking the current owner of an already-claimed
// forward to bring its worker back up.
//
// Neither of the existing channels can express "restart what already exists":
// SubmitTask is guarded by HasActiveTask and the entry is still present, so it
// dedupes; and the claim path drops any task for a forward that already has an
// owner (its defence against duplicate-claim collisions). Re-enabling a stopped
// forward therefore needs its own task kind, which the owner applies without
// re-running the claim bookkeeping.
func (e *Engine) SubmitRestart(local store.Local, remote store.Remote, link store.Link) *Task {
// The task must not carry the stop: the whole point is to re-enable.
link.Disabled = false
t := &Task{
ID: e.state.NextTaskID(),
Local: local,
Remote: remote,
Link: link,
Created: time.Now().Unix(),
Restart: true,
}
if e.state.PendingTasks == nil {
e.state.PendingTasks = map[string]*Task{}
}
e.state.PendingTasks[t.ID] = t
return t
}
// HasActiveTask reports whether a forward with this key is genuinely in flight
// or running: a non-revocation pending task, or a topology entry that is not
// marked disabled.
//
// It is deliberately narrower than HasTask. Under mark-don't-remove a stopped
// forward keeps its topology entry (that entry is what carries the flag between
// nodes), so "an entry exists" no longer means "this forward is active".
// SubmitTask must use THIS predicate, otherwise starting a stopped forward —
// startForward() publishes the enable and then calls SubmitTask — would be
// swallowed by its own idempotency guard.
func (e *Engine) HasActiveTask(local, remote string, port int) bool {
for _, t := range e.state.PendingList() {
if t.Revoke {
continue // a revocation is the opposite of an active forward
}
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
return true
}
}
for _, t := range e.state.TopologyList() {
if t.Link.Disabled {
continue // stopped; the entry is only a flag carrier
}
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
return true
}
}
return false
}
// HasTask reports whether ANY record for this forward exists — a pending task
// (including an in-flight revocation) or a topology entry (including one marked
// disabled). This is the "have we already handled this key" question, used to
// avoid re-issuing work on repeated canvas saves. For "is it actually running"
// use HasActiveTask instead.
func (e *Engine) HasTask(local, remote string, port int) bool { func (e *Engine) HasTask(local, remote string, port int) bool {
for _, t := range e.state.PendingList() { for _, t := range e.state.PendingList() {
if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port { if t.Local.Name == local && t.Remote.Name == remote && t.Link.RemotePort == port {
@ -828,6 +1016,78 @@ func (e *Engine) UpdateTopologyGroup(local, remote string, port int, group strin
return e.state.UpdateTopologyGroup(local, remote, port, group) return e.state.UpdateTopologyGroup(local, remote, port, group)
} }
// UpdateTopologyDisabled marks a forward enabled/disabled in the topology so
// the decision rides the next token cycle to every member. Paired with
// TopologyDisabled, which the adoption hook uses to learn peers' decisions.
//
// The decision is also remembered locally until the ring confirms it, because
// adoption replaces the whole state: a token captured before this call would
// otherwise wash the decision out on the next cycle and it would never reach
// the other members. See the localDisabled field.
func (e *Engine) UpdateTopologyDisabled(local, remote string, port int, disabled bool) bool {
if e.localDisabled == nil {
e.localDisabled = map[string]bool{}
}
e.localDisabled[disableKey(local, remote, port)] = disabled
return e.state.UpdateTopologyDisabled(local, remote, port, disabled)
}
// TopologyDisabled reports the cluster's view of a forward's disabled flag.
// The bool is false when the forward has no topology entry (nothing to learn).
//
// A pending local decision takes precedence over the adopted state: it has not
// had a chance to reach the other members yet, and reporting the adopted value
// would make the local store (and the status page) disagree with the user's
// most recent action.
func (e *Engine) TopologyDisabled(local, remote string, port int) (bool, bool) {
if d, pending := e.localDisabled[disableKey(local, remote, port)]; pending {
return d, true
}
return e.state.TopologyDisabled(local, remote, port)
}
// reconcileLocalDisabled re-asserts this node's not-yet-confirmed disabled
// decisions onto the freshly adopted state. Called right after e.state is
// replaced, and paired with reconcileTopologySync (which pushes those decisions
// out to the ring).
//
// A decision is dropped once the adopted state reports the SAME value: at that
// point every member agrees and there is nothing left to protect. Entries whose
// forward no longer exists in the topology are dropped too, so the map tracks
// only live disagreements.
func (e *Engine) reconcileLocalDisabled() {
if len(e.localDisabled) == 0 {
return
}
for id, te := range e.state.Topology {
_ = id
k := disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)
want, pending := e.localDisabled[k]
if !pending {
continue
}
if te.Link.Disabled == want {
// The ring caught up (or agreed independently): stop tracking.
delete(e.localDisabled, k)
continue
}
// Still divergent: keep asserting our decision onto the adopted state.
te.Link.Disabled = want
te.Active = !want
}
// Forget decisions for forwards that left the topology entirely — otherwise
// a long-lived cluster would accumulate dead keys.
live := make(map[string]struct{}, len(e.state.Topology))
for _, te := range e.state.Topology {
live[disableKey(te.Local.Name, te.Remote.Name, te.Link.RemotePort)] = struct{}{}
}
for k := range e.localDisabled {
if _, ok := live[k]; !ok {
delete(e.localDisabled, k)
}
}
}
// IsLeader reports whether this node is the current ring leader. // IsLeader reports whether this node is the current ring leader.
func (e *Engine) IsLeader() bool { return e.state.LeaderID == e.ID } func (e *Engine) IsLeader() bool { return e.state.LeaderID == e.ID }
@ -850,6 +1110,14 @@ func (e *Engine) SetPeerPersist(fn func(peersJSON string) error) { e.peerPersist
// HTTP handlers are overwritten by the next e.state = tk.State. // HTTP handlers are overwritten by the next e.state = tk.State.
func (e *Engine) SetTopologySync(fn func()) { e.topologySync = fn } func (e *Engine) SetTopologySync(fn func()) { e.topologySync = fn }
// disableKey builds the map key for a forward's disabled decision: the natural
// triple (local, remote, remotePort), which is how every other part of the code
// identifies a forward. A NUL separator keeps it unambiguous for names that
// could otherwise collide across the boundaries.
func disableKey(local, remote string, port int) string {
return local + "\x00" + remote + "\x00" + strconv.Itoa(port)
}
// persistPeers extracts all alive peers (addr + nodeKey, excluding self) // persistPeers extracts all alive peers (addr + nodeKey, excluding self)
// from the current ring state and persists them via the peerPersist callback. // from the current ring state and persists them via the peerPersist callback.
// Called on every token cycle (OnToken) and on AdoptState so a crashed node // Called on every token cycle (OnToken) and on AdoptState so a crashed node

View File

@ -11,6 +11,7 @@ type fakeHandler struct {
load Load load Load
claim func(ctx context.Context, tk *Task) error claim func(ctx context.Context, tk *Task) error
revoke func(ctx context.Context, tk *Task) error revoke func(ctx context.Context, tk *Task) error
restart func(ctx context.Context, tk *Task) error
} }
func (h *fakeHandler) Revoke(ctx context.Context, tk *Task) error { func (h *fakeHandler) Revoke(ctx context.Context, tk *Task) error {
@ -20,6 +21,13 @@ func (h *fakeHandler) Revoke(ctx context.Context, tk *Task) error {
return nil return nil
} }
func (h *fakeHandler) Restart(ctx context.Context, tk *Task) error {
if h.restart != nil {
return h.restart(ctx, tk)
}
return nil
}
func (h *fakeHandler) RuntimeLoad() Load { return h.load } func (h *fakeHandler) RuntimeLoad() Load { return h.load }
func (h *fakeHandler) Claim(ctx context.Context, tk *Task) error { func (h *fakeHandler) Claim(ctx context.Context, tk *Task) error {
if h.claim != nil { if h.claim != nil {

View File

@ -17,6 +17,10 @@ type AppHandler struct {
ClaimFn func(ctx context.Context, tk *Task) error ClaimFn func(ctx context.Context, tk *Task) error
// RevokeFn cancels the forward on this node (stop worker + drop from store). // RevokeFn cancels the forward on this node (stop worker + drop from store).
RevokeFn func(ctx context.Context, tk *Task) error RevokeFn func(ctx context.Context, tk *Task) error
// RestartFn re-enables an already-claimed forward on this node (clear the
// stopped flag + bring the worker back). Required because a stop keeps the
// topology entry, so the normal claim path will not re-create it.
RestartFn func(ctx context.Context, tk *Task) error
} }
// RuntimeLoad implements Handler. // RuntimeLoad implements Handler.
@ -44,6 +48,14 @@ func (a *AppHandler) Revoke(ctx context.Context, tk *Task) error {
return a.RevokeFn(ctx, tk) return a.RevokeFn(ctx, tk)
} }
// Restart implements Handler.
func (a *AppHandler) Restart(ctx context.Context, tk *Task) error {
if a.RestartFn == nil {
return nil
}
return a.RestartFn(ctx, tk)
}
// SampleMemLoad returns a cheap memory-usage percentage (0..100). // SampleMemLoad returns a cheap memory-usage percentage (0..100).
func SampleMemLoad() float64 { func SampleMemLoad() float64 {
var m runtime.MemStats var m runtime.MemStats

View File

@ -128,7 +128,7 @@ func (e *Engine) forwardToNext(ctx context.Context, tk *Token) error {
if e.send == nil { if e.send == nil {
return nil return nil
} }
log.Printf("ring[%s] forward cycle=%d to %s", e.ID, tk.Cycle, next) debugToken("ring[%s] forward cycle=%d to %s", e.ID, tk.Cycle, next)
err := e.send(ctx, next, tk) err := e.send(ctx, next, tk)
if err == nil { if err == nil {
if e.state.LeaderID == e.ID { if e.state.LeaderID == e.ID {

View File

@ -20,6 +20,8 @@ import (
// LogKind enumerates operation-log entry types. // LogKind enumerates operation-log entry types.
const ( const (
LogForwardAdd = "forward.add" LogForwardAdd = "forward.add"
LogForwardStop = "forward.stop"
LogForwardStart = "forward.start"
LogForwardRemove = "forward.remove" LogForwardRemove = "forward.remove"
LogNodeJoin = "node.join" LogNodeJoin = "node.join"
LogNodeLeave = "node.leave" LogNodeLeave = "node.leave"
@ -140,7 +142,7 @@ func DetailOf(e LogEntry) string {
} }
str := func(key string) string { s, _ := d[key].(string); return s } str := func(key string) string { s, _ := d[key].(string); return s }
switch e.Kind { switch e.Kind {
case LogForwardAdd, LogForwardRemove: case LogForwardAdd, LogForwardStop, LogForwardStart, LogForwardRemove:
s := str("local") + " → " + str("remote") s := str("local") + " → " + str("remote")
if id := str("taskId"); id != "" { if id := str("taskId"); id != "" {
if len(id) > 8 { if len(id) > 8 {

View File

@ -2,14 +2,22 @@ package cluster
import ( import (
"context" "context"
"encoding/json"
"testing" "testing"
"webui4frpc/internal/store" "webui4frpc/internal/store"
) )
// TestRevokeTaskRemovesTopology: a revoke task published to the ring removes // TestRevokeTaskMarksTopologyDisabled: a stop delivered through the ring marks
// the forward from topology, calls Handler.Revoke, and logs forward.remove. // the topology entry disabled instead of deleting it, calls Handler.Revoke, and
func TestRevokeTaskRemovesTopology(t *testing.T) { // logs forward.stop.
//
// The entry is deliberately KEPT: TopoEntry.Link is the carrier that makes the
// stop travel, and State rides the token every cycle. Keeping it is what lets a
// stop reach a member that was offline when the stop was issued — deleting the
// entry left the flag nowhere to live, so a stop converged only if the owning
// node happened to be online for a one-shot revoke task.
func TestRevokeTaskMarksTopologyDisabled(t *testing.T) {
revoked := false revoked := false
eng := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil, eng := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil,
&fakeHandler{load: Load{MemPct: 5, NetPct: 5}, &fakeHandler{load: Load{MemPct: 5, NetPct: 5},
@ -26,33 +34,92 @@ func TestRevokeTaskRemovesTopology(t *testing.T) {
if len(eng.state.TopologyList()) != 1 { if len(eng.state.TopologyList()) != 1 {
t.Fatalf("topology after claim = %+v", eng.state.TopologyList()) t.Fatalf("topology after claim = %+v", eng.state.TopologyList())
} }
if d, _ := eng.state.TopologyDisabled("web", "frps1", 18081); d {
t.Fatal("a freshly claimed forward must not start out disabled")
}
// publish a revoke task pointing at the same forward // publish a revoke task pointing at the same forward
eng.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081}) eng.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state}); err != nil { if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state}); err != nil {
t.Fatal(err) t.Fatal(err)
} }
if len(eng.state.TopologyList()) != 0 {
t.Fatalf("topology after revoke = %+v", eng.state.TopologyList()) // The entry must SURVIVE, marked disabled, so the flag keeps travelling.
if len(eng.state.TopologyList()) != 1 {
t.Fatalf("topology entry was removed; the disabled flag would have no carrier: %+v",
eng.state.TopologyList())
}
disabled, known := eng.state.TopologyDisabled("web", "frps1", 18081)
if !known {
t.Fatal("topology entry missing after revoke")
}
if !disabled {
t.Fatal("revoke did not mark the topology entry disabled")
}
// Active must follow the flag, since OfflineReassign only re-queues Active
// entries — a stopped forward must not be resurrected by a node departure.
if eng.state.TopologyList()[0].Active {
t.Fatal("a stopped forward must not remain Active (OfflineReassign would re-queue it)")
} }
if !revoked { if !revoked {
t.Fatal("Handler.Revoke was not called") t.Fatal("Handler.Revoke was not called")
} }
// log should contain forward.remove // log should contain forward.stop (not forward.remove — nothing was removed)
var sawRemove bool var sawStop bool
for _, e := range eng.Log.Snapshot() { for _, e := range eng.Log.Snapshot() {
if e.Kind == LogForwardRemove { if e.Kind == LogForwardStop {
sawRemove = true sawStop = true
} }
} }
if !sawRemove { if !sawStop {
t.Fatalf("log missing forward.remove: %+v", eng.Log.Snapshot()) t.Fatalf("log missing forward.stop: %+v", eng.Log.Snapshot())
}
}
// TestStoppedTopologyEntrySurvivesAdoption is the property that makes the whole
// mark-don't-remove design work: because the entry persists, a node that adopts
// ring state learns the stop. This is the offline-member case a one-shot revoke
// task could never cover.
func TestStoppedTopologyEntrySurvivesAdoption(t *testing.T) {
src := newTestEngine("n1", true)
src.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := src.OnToken(context.Background(), &Token{Cycle: 1, State: src.state}); err != nil {
t.Fatal(err)
}
src.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := src.OnToken(context.Background(), &Token{Cycle: 2, State: src.state}); err != nil {
t.Fatal(err)
}
// A different node adopts the ring state (what e.state = tk.State does).
peer := newTestEngine("n2", false)
peer.AdoptState(src.state)
disabled, known := peer.TopologyDisabled("web", "frps1", 18081)
if !known {
t.Fatal("peer did not receive the topology entry for the stopped forward")
}
if !disabled {
t.Fatal("peer adopted the entry but no longer sees it as disabled — the stop did not propagate")
} }
} }
// TestRevokeIdempotent: revoking an already-missing forward does not error. // TestRevokeIdempotent: revoking an already-missing forward does not error.
//
// It must ALSO still call Handler.Revoke. The handler is what actually stops
// the per-forward frpc worker; the topology entry is only bookkeeping. Gating
// the handler on RemoveTopology()'s return value (as this test used to permit)
// turned every revoke of an already-absent forward into a silent no-op — the
// worker kept running, which is how a forward the user had stopped kept
// dialling a local service that was intentionally down, forever.
func TestRevokeIdempotent(t *testing.T) { func TestRevokeIdempotent(t *testing.T) {
eng := newTestEngine("n1", true) revoked := 0
eng := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil,
&fakeHandler{load: Load{MemPct: 5, NetPct: 5},
revoke: func(ctx context.Context, tk *Task) error { revoked++; return nil }},
func(ctx context.Context, next string, tk *Token) error { return nil },
"n1:7500", true, "")
eng.state.AddRevoke(store.Local{Name: "ghost"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 1}) eng.state.AddRevoke(store.Local{Name: "ghost"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 1})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil { if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil {
t.Fatalf("revoke missing: %v", err) t.Fatalf("revoke missing: %v", err)
@ -61,4 +128,318 @@ func TestRevokeIdempotent(t *testing.T) {
if len(eng.state.TopologyList()) != 0 { if len(eng.state.TopologyList()) != 0 {
t.Fatal("should be empty") t.Fatal("should be empty")
} }
// The stop side-effect must have happened even though there was no entry.
if revoked != 1 {
t.Fatalf("Handler.Revoke called %d times, want 1 — a revoke with no topology entry "+
"must still stop the worker, otherwise stopped forwards keep running", revoked)
}
}
// TestStoppedForwardNotRequeuedOnNodeDeparture pins why marking (rather than
// deleting) is safe: OfflineReassign only re-queues ACTIVE entries, so a
// stopped forward is not silently handed to another node when its owner goes
// away. Deleting the entry, or leaving it Active, would both resurrect it.
func TestStoppedForwardNotRequeuedOnNodeDeparture(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil {
t.Fatal(err)
}
eng.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state}); err != nil {
t.Fatal(err)
}
owner := eng.state.TopologyList()[0].OwnerID
eng.state.OfflineReassign(owner)
if n := len(eng.state.PendingList()); n != 0 {
t.Fatalf("a stopped forward was re-queued on the owner's departure (pending=%d): %+v",
n, eng.state.PendingList())
}
}
// TestSubmitTaskNotBlockedByStoppedEntry is the regression test for the bug the
// mark-don't-remove change introduced: SubmitTask's idempotency guard used
// HasTask, which matches a disabled topology entry. Since a stop now KEEPS the
// entry, any submission for that forward while it is still stopped is swallowed
// by its own guard — so a stopped forward can never be started again.
//
// The probe deliberately submits WITHOUT re-enabling first. Re-enabling would
// clear the flag and make HasTask and HasActiveTask agree, hiding the defect;
// the guard has to be exercised at the moment the entry is still stopped.
func TestSubmitTaskNotBlockedByStoppedEntry(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil {
t.Fatal(err)
}
// stop it -> entry is kept, marked disabled
eng.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state}); err != nil {
t.Fatal(err)
}
if d, _ := eng.state.TopologyDisabled("web", "frps1", 18081); !d {
t.Fatal("precondition: forward should be disabled")
}
// The stopped entry IS findable by the broad predicate — which is exactly why
// the narrow one has to exist.
if !eng.HasTask("web", "frps1", 18081) {
t.Fatal("precondition: the stopped entry should still be findable by HasTask")
}
if eng.HasActiveTask("web", "frps1", 18081) {
t.Fatal("precondition: a stopped forward must NOT count as active")
}
// The regression: a submission must be published for a stopped forward.
// With a HasTask guard this returns nil and the forward is stuck forever.
tk := eng.SubmitTask(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081})
if tk == nil {
t.Fatal("SubmitTask was swallowed by the stopped topology entry — " +
"a stopped forward could never be started again")
}
if tk.Link.RemotePort != 18081 {
t.Fatalf("unexpected task: %+v", tk)
}
}
// TestSubmitTaskStillDedupesActiveForward is the other half of the guard: the
// narrower predicate must not turn SubmitTask into a duplicate-task generator.
func TestSubmitTaskStillDedupesActiveForward(t *testing.T) {
eng := newTestEngine("n1", true)
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil {
t.Fatal(err)
}
if !eng.HasActiveTask("web", "frps1", 18081) {
t.Fatal("precondition: a claimed forward must count as active")
}
if tk := eng.SubmitTask(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081}); tk != nil {
t.Fatalf("SubmitTask must stay idempotent for an active forward, got %+v", tk)
}
}
// TestStoppedForwardNotRequeuedOnNodeDeparture pins why marking (rather than
// deleting) is safe: OfflineReassign only re-queues ACTIVE entries, so a
// stopped forward is not silently handed to another node when its owner goes
// TestAddTopologyRespectsDisabledFlag: a claim that carries a stopped forward
// must not mark the new entry Active, or it would come back to life.
func TestAddTopologyRespectsDisabledFlag(t *testing.T) {
s := &State{Topology: map[string]*TopoEntry{}}
e := s.AddTopology(&Task{
ID: "t1", Local: store.Local{Name: "web"}, Remote: store.Remote{Name: "frps1"},
Link: store.Link{RemotePort: 18081, Disabled: true},
}, "n1")
if e.Active {
t.Fatal("AddTopology forced Active=true for a disabled claim — the forward would resurrect")
}
if !e.Link.Disabled {
t.Fatal("AddTopology dropped the disabled flag from the link")
}
}
// TestRestartTaskBypassesDuplicateClaimGuard is the engine half of the
// re-enable path: a restart task for a forward that still HAS a topology entry
// must reach Handler.Restart.
//
// This is precisely what the duplicate-claim guard refuses — it exists to stop
// a stale task from spawning a second worker for an owned forward, and a restart
// is indistinguishable from that unless it is an explicit task kind. Marking
// the entry stopped (instead of deleting it) is what made this necessary: the
// entry the guard keys on now outlives a stop.
func TestRestartTaskBypassesDuplicateClaimGuard(t *testing.T) {
restarts := 0
claims := 0
eng := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil,
&fakeHandler{load: Load{MemPct: 5, NetPct: 5},
claim: func(ctx context.Context, tk *Task) error { claims++; return nil },
restart: func(ctx context.Context, tk *Task) error { restarts++; return nil }},
func(ctx context.Context, next string, tk *Token) error { return nil },
"n1:7500", true, "")
// Establish + stop: entry survives, marked disabled.
eng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 1, State: eng.state}); err != nil {
t.Fatal(err)
}
eng.state.AddRevoke(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 2, State: eng.state}); err != nil {
t.Fatal(err)
}
if d, _ := eng.state.TopologyDisabled("web", "frps1", 18081); !d {
t.Fatal("precondition: should be disabled")
}
entryCount := len(eng.state.TopologyList())
// Start: publish a restart and run a cycle.
eng.state.UpdateTopologyDisabled("web", "frps1", 18081, false)
eng.SubmitRestart(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081})
if _, err := eng.OnToken(context.Background(), &Token{Cycle: 3, State: eng.state}); err != nil {
t.Fatal(err)
}
if restarts != 1 {
t.Fatalf("Handler.Restart called %d times, want 1 — the restart task was swallowed "+
"by the duplicate-claim guard, so the forward would never come back", restarts)
}
if claims != 1 {
t.Fatalf("Handler.Claim called %d times, want 1 (the original claim only); "+
"a restart must not re-run the claim bookkeeping", claims)
}
// Re-enabling must not create a second entry.
if n := len(eng.state.TopologyList()); n != entryCount {
t.Fatalf("restart created a duplicate topology entry: %d -> %d", entryCount, n)
}
if d, _ := eng.state.TopologyDisabled("web", "frps1", 18081); d {
t.Fatal("entry is still disabled after the restart task was applied")
}
var sawStart bool
for _, e := range eng.Log.Snapshot() {
if e.Kind == LogForwardStart {
sawStart = true
}
}
if !sawStart {
t.Fatalf("log missing forward.start: %+v", eng.Log.Snapshot())
}
}
// TestRestartFlagSurvivesTokenSerialization: the restart task travels to the
// owner inside the token, so the flag must round-trip through JSON. Without the
// struct tag it would silently deserialize as false and the owner would treat
// it as an ordinary claim — which the duplicate guard then discards.
func TestRestartFlagSurvivesTokenSerialization(t *testing.T) {
eng := newTestEngine("n1", true)
eng.SubmitRestart(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081})
blob, err := json.Marshal(&Token{Cycle: 7, State: eng.state})
if err != nil {
t.Fatal(err)
}
var back Token
if err := json.Unmarshal(blob, &back); err != nil {
t.Fatal(err)
}
var found *Task
for _, tk := range back.State.PendingList() {
if tk.Local.Name == "web" && tk.Link.RemotePort == 18081 {
found = tk
}
}
if found == nil {
t.Fatal("restart task did not survive token serialization")
}
if !found.Restart {
t.Fatal("the restart flag was lost in JSON — the owner would see a plain claim " +
"and the duplicate guard would discard it")
}
// Revoke and Restart must stay distinguishable.
if found.Revoke {
t.Fatal("a restart task must not also read as a revocation")
}
}
// TestRestartTaskReachesNonLocalOwner.
func TestRestartOnlyAppliedByOwner(t *testing.T) {
restarts := 0
h := &fakeHandler{load: Load{MemPct: 5, NetPct: 5},
restart: func(ctx context.Context, tk *Task) error { restarts++; return nil }}
// Owner node claims the forward.
owner := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil, h,
func(ctx context.Context, next string, tk *Token) error { return nil }, "n1:7500", true, "")
owner.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := owner.OnToken(context.Background(), &Token{Cycle: 1, State: owner.state, SentAt: 1}); err != nil {
t.Fatal(err)
}
if len(owner.state.TopologyList()) != 1 {
t.Fatalf("precondition: owner should hold the entry, got %+v", owner.state.TopologyList())
}
// A peer adopts the same ring state but does not own the forward.
peer := NewEngine("n2", "n2:7500", "u", "p", "0.71.0", nil, h,
func(ctx context.Context, next string, tk *Token) error { return nil }, "n2:7500", false, "")
peer.AdoptState(owner.state)
tk := peer.SubmitRestart(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081})
if tk == nil {
t.Fatal("SubmitRestart returned nil")
}
// The peer processes one token: it must NOT apply the restart.
if _, err := peer.OnToken(context.Background(), &Token{Cycle: 2, State: peer.state, SentAt: 2}); err != nil {
t.Fatal(err)
}
if restarts != 0 {
t.Fatalf("a NON-owner applied the restart %d time(s) — it would spawn an orphaned "+
"worker for a forward owned by somebody else", restarts)
}
}
// TestRestartTaskReachesNonLocalOwner: a restart published on a node that is NOT
// the owner must survive the token round-trip and be applied by the owner.
//
// The naive "if not owner then continue" implementation CONSUMED the task
// (ClaimPending had already removed it), so the owner never received it and the
// forward stayed stopped with nothing logged as an error. Deferral must
// re-queue it.
//
// Each OnToken is fed a FRESH token (as the real ring does — the successor's
// OnToken receives the state the predecessor returned), so the deferral is
// exercised once per hop rather than re-processing one snapshot.
func TestRestartTaskReachesNonLocalOwner(t *testing.T) {
ownerApplied, submitterApplied := 0, 0
ownerEng := NewEngine("n2", "n2:7500", "u", "p", "0.71.0", nil,
&fakeHandler{load: Load{MemPct: 5, NetPct: 5},
restart: func(ctx context.Context, tk *Task) error { ownerApplied++; return nil }},
func(ctx context.Context, next string, tk *Token) error { return nil }, "n2:7500", false, "")
ownerEng.state.AddPending(store.Local{Name: "web"}, store.Remote{Name: "frps1"}, store.Link{RemotePort: 18081})
if _, err := ownerEng.OnToken(context.Background(), &Token{Cycle: 1, State: ownerEng.state}); err != nil {
t.Fatal(err)
}
if len(ownerEng.state.TopologyList()) != 1 {
t.Fatalf("precondition: n2 should own it, got %+v", ownerEng.state.TopologyList())
}
subEng := NewEngine("n1", "n1:7500", "u", "p", "0.71.0", nil,
&fakeHandler{load: Load{MemPct: 5, NetPct: 5},
restart: func(ctx context.Context, tk *Task) error { submitterApplied++; return nil }},
func(ctx context.Context, next string, tk *Token) error { return nil }, "n1:7500", true, "")
subEng.AdoptState(ownerEng.state)
subEng.SubmitRestart(store.Local{Name: "web"}, store.Remote{Name: "frps1"},
store.Link{RemotePort: 18081})
// One hop through the non-owner: it must defer, not apply and not drop.
out, err := subEng.OnToken(context.Background(), &Token{Cycle: 2, State: subEng.state, SentAt: 1})
if err != nil {
t.Fatal(err)
}
if submitterApplied != 0 {
t.Fatal("the non-owner applied a restart for a forward it does not own")
}
var carried *Task
for _, tk := range out.State.PendingList() {
if tk.Restart && tk.Local.Name == "web" {
carried = tk
}
}
if carried == nil {
t.Fatal("the non-owner CONSUMED the restart task; the owner would never receive it")
}
// The owner applies the carried task.
if _, err := ownerEng.OnToken(context.Background(), &Token{Cycle: 3, State: out.State, SentAt: 2}); err != nil {
t.Fatal(err)
}
if ownerApplied != 1 {
t.Fatalf("the owner applied the restart %d times, want 1", ownerApplied)
}
} }

View File

@ -8,7 +8,6 @@ import (
"context" "context"
"encoding/json" "encoding/json"
"fmt" "fmt"
"log"
"net/http" "net/http"
"time" "time"
) )
@ -52,11 +51,16 @@ func (t *TokenTransport) SendTo(getAddr func(nodeID string) string) func(ctx con
start := time.Now() start := time.Now()
resp, err := cli.Do(req) resp, err := cli.Do(req)
if err != nil { if err != nil {
log.Printf("token-send %s: size=%dB elapsed=%v err=%v", addr, len(body), time.Since(start).Round(time.Millisecond), err) // NOT demoted: a failed send is the "neighbor offline" signal the
// ring's fault paths are diagnosed from.
logf("token-send %s: size=%dB elapsed=%v err=%v", addr, len(body), time.Since(start).Round(time.Millisecond), err)
return err return err
} }
defer resp.Body.Close() defer resp.Body.Close()
log.Printf("token-send %s: size=%dB elapsed=%v -> %d", addr, len(body), time.Since(start).Round(time.Millisecond), resp.StatusCode) // Success path is a per-round heartbeat (size/elapsed/200 repeat
// verbatim every cycle) — debug only. The status-code check below stays
// loud on purpose, so a non-2xx still shows up without the debug flag.
debugToken("token-send %s: size=%dB elapsed=%v -> %d", addr, len(body), time.Since(start).Round(time.Millisecond), resp.StatusCode)
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return fmt.Errorf("token POST %s -> HTTP %d", url, resp.StatusCode) return fmt.Errorf("token POST %s -> HTTP %d", url, resp.StatusCode)
} }

View File

@ -0,0 +1,204 @@
package httpapi
import (
"bytes"
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"path/filepath"
"testing"
"webui4frpc/internal/cluster"
"webui4frpc/internal/process"
"webui4frpc/internal/store"
)
// newRingTestHandler builds a Handler WITH a ring engine attached, so the
// paths that publish tasks into the token actually execute. newTestHandler
// leaves Ring nil, which silently skips them — a test built on it can pass
// while the publish side is completely broken.
func newRingTestHandler(t *testing.T) (*Handler, *cluster.Engine, *httptest.Server) {
t.Helper()
dir := t.TempDir()
st, err := store.New(filepath.Join(dir, "test.db"))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = st.Close() })
pm := process.NewManager(process.Options{
ConfigsDir: filepath.Join(dir, "configs"),
LogsDir: filepath.Join(dir, "logs"),
BinaryPath: func() string { return "" },
Render: func(string) ([]byte, error) { return []byte(`{}`), nil },
AutoRestart: func(string) bool { return false },
RestartInterval: func() int { return 5 },
})
ring := cluster.NewEngine("n1", "n1:7500", "u", "p", "0.1.0", nil,
&cluster.AppHandler{},
func(ctx context.Context, next string, tk *cluster.Token) error { return nil },
"n1:7500", true, "")
h := &Handler{Store: st, Process: pm, WorkDir: dir, User: "admin", Password: "pw", Ring: ring}
mux, err := NewServeMux(h)
if err != nil {
t.Fatal(err)
}
ts := httptest.NewServer(mux)
t.Cleanup(ts.Close)
return h, ring, ts
}
func saveCanvas(t *testing.T, srv *httptest.Server, body string) {
t.Helper()
req, _ := http.NewRequest(http.MethodPut, srv.URL+"/api/manager/canvas", bytes.NewBufferString(body))
req.SetBasicAuth("admin", "pw")
resp, err := srv.Client().Do(req)
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("canvas save status = %d", resp.StatusCode)
}
}
// TestStopForwardPublishesDisabledFlagInRevokeTask is the regression test for
// the actual defect.
//
// stopForward() read the link, called SetLinkDisabled(true), and then handed
// the STALE copy (disabled=false) to RevokeTask. The revoke travels to the
// node that OWNS the forward, and that node's Claim/Revoke path keys off the
// flag — so a stale false meant:
// - the owner could not tell the stop was deliberate, and
// - nothing retired the topology entry,
//
// so the next restart's reconcile re-claimed the forward and spawned a worker
// for something the user had stopped (seen live: endless connection-refused
// against an intentionally-down service).
//
// This asserts the flag ON THE PUBLISHED TASK, which is the value that was
// actually wrong. It cannot be satisfied by the store write alone.
func TestStopForwardPublishesDisabledFlagInRevokeTask(t *testing.T) {
_, ring, ts := newRingTestHandler(t)
saveCanvas(t, ts, `{
"locals": [{"name":"svc","ip":"127.0.0.1","port":59999,"protocol":"tcp"}],
"remotes": [{"name":"srv-a","ip":"1.2.3.4","port":7000,"enabled":true}],
"links": [{"local":"svc","remote":"srv-a","remotePort":45999}]
}`)
// Stop the forward over the API.
b, _ := json.Marshal(stopForwardReq{"svc", "srv-a", 45999})
req, _ := http.NewRequest(http.MethodPost, ts.URL+"/api/manager/forwards/stop", bytes.NewReader(b))
req.SetBasicAuth("admin", "pw")
resp, err := ts.Client().Do(req)
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("stop status = %d", resp.StatusCode)
}
// Find the revoke task that was published into the token.
var revoke *cluster.Task
for _, tk := range ring.State().PendingList() {
if tk.Revoke && tk.Local.Name == "svc" && tk.Link.RemotePort == 45999 {
revoke = tk
break
}
}
if revoke == nil {
t.Fatal("stop did not publish a revoke task for the forward")
}
if !revoke.Link.Disabled {
t.Fatal("the published revoke task carries disabled=false — the owner node cannot tell " +
"this stop was deliberate, which is the bug that let stopped forwards resurrect")
}
}
// TestStopThenStartPublishesRestartTask guards the failure mode that
// mark-don't-remove introduces, and that only shows up on the WIRE.
//
// A stop keeps the topology entry (it is the carrier for the flag), so on
// re-enable:
// - SubmitTask dedupes, because the entry exists;
// - the claim path discards any task for a forward that already has an owner.
//
// Both channels therefore refuse, no task is published, the owner never learns
// about the re-enable, and the forward stays stopped forever — a forward the
// user can stop but never restart.
//
// Asserting the task is PUBLISHED is the point: asserting only that the flag
// cleared would pass while nothing reached the owner. (The earlier version of
// this test did exactly that and stayed green against the broken code.)
func TestStopThenStartPublishesRestartTask(t *testing.T) {
_, ring, ts := newRingTestHandler(t)
saveCanvas(t, ts, `{
"locals": [{"name":"svc","ip":"127.0.0.1","port":59999,"protocol":"tcp"}],
"remotes": [{"name":"srv-a","ip":"1.2.3.4","port":7000,"enabled":true}],
"links": [{"local":"svc","remote":"srv-a","remotePort":45999}]
}`)
// Claim it so a topology entry exists (that is what makes the two normal
// channels refuse later).
if _, err := ring.OnToken(context.Background(), &cluster.Token{Cycle: 1, State: *ring.State()}); err != nil {
t.Fatal(err)
}
if !ring.HasActiveTask("svc", "srv-a", 45999) {
t.Fatalf("precondition: forward should be active, topo=%+v", ring.State().TopologyList())
}
forwards := func(action string) {
t.Helper()
b, _ := json.Marshal(stopForwardReq{"svc", "srv-a", 45999})
req, _ := http.NewRequest(http.MethodPost, ts.URL+"/api/manager/forwards/"+action, bytes.NewReader(b))
req.SetBasicAuth("admin", "pw")
resp, err := ts.Client().Do(req)
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("%s status = %d", action, resp.StatusCode)
}
}
// --- stop: entry kept, marked disabled -------------------------------
forwards("stop")
if d, known := ring.TopologyDisabled("svc", "srv-a", 45999); !known || !d {
t.Fatalf("stop did not mark the topology entry disabled (known=%v disabled=%v)", known, d)
}
if ring.HasActiveTask("svc", "srv-a", 45999) {
t.Fatal("a stopped forward must not count as active")
}
// --- start: a task MUST reach the owner ------------------------------
forwards("start")
if d, _ := ring.TopologyDisabled("svc", "srv-a", 45999); d {
t.Fatal("start did not clear the disabled flag")
}
if !ring.HasActiveTask("svc", "srv-a", 45999) {
t.Fatal("after start the forward must be active again")
}
// The actual regression: something has to be queued for the owner. The
// re-enable must be published as a restart task, since a plain creation task
// would be deduped or discarded.
var restart *cluster.Task
for _, tk := range ring.State().PendingList() {
if tk.Restart && tk.Local.Name == "svc" && tk.Link.RemotePort == 45999 {
restart = tk
break
}
}
if restart == nil {
t.Fatalf("start published no restart task — the owner would never bring the "+
"worker back, leaving the forward permanently stopped (pending=%+v)",
ring.State().PendingList())
}
if restart.Link.Disabled {
t.Fatal("the restart task carries disabled=true, so the owner would re-apply the stop")
}
}

View File

@ -0,0 +1,141 @@
package httpapi
import (
"bytes"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
)
// stopForwardReq mirrors the forwards start/stop request body.
type stopForwardReq struct {
Local string `json:"local"`
Remote string `json:"remote"`
RemotePort int `json:"remotePort"`
}
func postForwards(t *testing.T, srv *httptest.Server, action string, body stopForwardReq) int {
t.Helper()
b, _ := json.Marshal(body)
req, _ := http.NewRequest(http.MethodPost, srv.URL+"/api/manager/forwards/"+action, bytes.NewReader(b))
req.SetBasicAuth("admin", "pw")
req.Header.Set("Content-Type", "application/json")
resp, err := srv.Client().Do(req)
if err != nil {
t.Fatal(err)
}
defer resp.Body.Close()
return resp.StatusCode
}
// TestStopForwardPersistsDisabledFlag is the regression test for the
// "stopped forwards resurrect on restart" bug.
//
// The original defect: stopForward() read the link, called
// SetLinkDisabled(true), but then handed the STALE (disabled=false) copy to
// RevokeTask — so the disabled flag never reached the node owning the forward,
// and nothing removed the topology entry. On the next restart the startup
// reconcile saw an owned forward with no worker and re-claimed it, spawning a
// worker for a forward the user had deliberately stopped (observed live:
// ~14k "proxy already exists" retries and endless connection-refused against a
// service that was intentionally down).
//
// The contract this pins: after a successful stop, the persisted link MUST be
// disabled — that flag is the single source of truth the claim path consults.
func TestStopForwardPersistsDisabledFlag(t *testing.T) {
h, ts := newTestHandler(t)
body := `{
"locals": [{"name":"svc","ip":"127.0.0.1","port":59999,"protocol":"tcp"}],
"remotes": [{"name":"srv-a","ip":"1.2.3.4","port":7000,"enabled":true}],
"links": [{"local":"svc","remote":"srv-a","remotePort":45999}]
}`
req, _ := http.NewRequest(http.MethodPut, ts.URL+"/api/manager/canvas", bytes.NewBufferString(body))
req.SetBasicAuth("admin", "pw")
resp, err := ts.Client().Do(req)
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("canvas save status = %d", resp.StatusCode)
}
// The forward starts enabled.
ln, found, err := h.Store.LinkByTriple("svc", "srv-a", 45999)
if err != nil || !found {
t.Fatalf("link not persisted: found=%v err=%v", found, err)
}
if ln.Disabled {
t.Fatal("a freshly saved forward must start enabled")
}
// Stop it.
if code := postForwards(t, ts, "stop", stopForwardReq{"svc", "srv-a", 45999}); code != http.StatusOK {
t.Fatalf("stop status = %d, want 200", code)
}
// Persisted flag must now be set — this is what the claim path reads.
ln, found, err = h.Store.LinkByTriple("svc", "srv-a", 45999)
if err != nil {
t.Fatal(err)
}
if !found {
t.Fatal("link vanished after stop; stop must be non-destructive")
}
if !ln.Disabled {
t.Fatal("stop did not persist disabled=true — the startup reconcile would resurrect this forward")
}
// Start must clear it again (the user-facing re-enable path).
if code := postForwards(t, ts, "start", stopForwardReq{"svc", "srv-a", 45999}); code != http.StatusOK {
t.Fatalf("start status = %d, want 200", code)
}
ln, _, err = h.Store.LinkByTriple("svc", "srv-a", 45999)
if err != nil {
t.Fatal(err)
}
if ln.Disabled {
t.Fatal("start did not clear disabled; a re-enabled forward would stay stopped")
}
}
// TestStopForwardIsNonDestructive pins the per-forward stop semantics the
// revoke path was specifically rewritten for: stopping one forward must not
// touch a sibling forward that shares the same local or remote.
func TestStopForwardIsNonDestructive(t *testing.T) {
h, ts := newTestHandler(t)
body := `{
"locals": [{"name":"svc","ip":"127.0.0.1","port":59999,"protocol":"tcp"}],
"remotes": [{"name":"srv-a","ip":"1.2.3.4","port":7000,"enabled":true}],
"links": [
{"local":"svc","remote":"srv-a","remotePort":45999},
{"local":"svc","remote":"srv-a","remotePort":46000}
]
}`
req, _ := http.NewRequest(http.MethodPut, ts.URL+"/api/manager/canvas", bytes.NewBufferString(body))
req.SetBasicAuth("admin", "pw")
resp, err := ts.Client().Do(req)
if err != nil {
t.Fatal(err)
}
resp.Body.Close()
if code := postForwards(t, ts, "stop", stopForwardReq{"svc", "srv-a", 45999}); code != http.StatusOK {
t.Fatalf("stop status = %d", code)
}
stopped, _, _ := h.Store.LinkByTriple("svc", "srv-a", 45999)
sibling, found, _ := h.Store.LinkByTriple("svc", "srv-a", 46000)
if !found {
t.Fatal("sibling forward was destroyed by stopping its neighbour")
}
if !stopped.Disabled {
t.Error("the stopped forward should be disabled")
}
if sibling.Disabled {
t.Error("the sibling forward must stay enabled — per-forward stop, not per-local/remote")
}
}

View File

@ -245,7 +245,12 @@ func (h *Handler) applyCanvas(w http.ResponseWriter, r *http.Request, canvas *ca
} }
if ln.Disabled { if ln.Disabled {
// Stopped on the forwards page: make sure it leaves the topology. // Stopped on the forwards page: make sure it leaves the topology.
if h.Ring.HasTask(ln.Local, ln.Remote, ln.RemotePort) { // Guard on an ACTIVE task — an entry that is already marked
// disabled has nothing left to revoke, and re-issuing a revocation
// for it on every canvas save would be pure noise. (HasTask, which
// also matches disabled entries, is the right predicate for the
// "already handled" question this branch is not asking.)
if h.Ring.HasActiveTask(ln.Local, ln.Remote, ln.RemotePort) {
h.Ring.RevokeTask(loc, rem, ln) h.Ring.RevokeTask(loc, rem, ln)
} }
continue continue

View File

@ -84,6 +84,25 @@ func (h *Handler) startForward(local, remote string, port int) error {
} }
return nil return nil
} }
if h.Ring != nil {
// Publish the enable into the topology so it rides the ring (a peer that
// still holds the stale "stopped" copy learns about it on adoption).
h.Ring.UpdateTopologyDisabled(local, remote, port, false)
}
// A cluster forward whose topology entry still exists needs its OWNER to
// bring the worker back, and neither of the normal channels can do it:
// - SubmitTask is idempotency-guarded, and the entry still exists (a stop
// keeps it as the flag's carrier), so the submission is deduped away;
// - the claim path drops any task for a forward that already has an
// owner, as a defence against duplicate-claim collisions.
// So a re-enable of an existing entry is published as a dedicated RESTART
// task, which the owner applies unconditionally. Without it a forward the
// user can stop but not restart — which is what marking-instead-of-removing
// would otherwise have produced.
if h.Ring != nil && h.Ring.HasTopologyEntry(local, remote, port) {
h.Ring.SubmitRestart(loc, rem, ln)
return nil
}
if h.Ring != nil { if h.Ring != nil {
h.Ring.SubmitTask(loc, rem, ln) h.Ring.SubmitTask(loc, rem, ln)
} }
@ -109,6 +128,24 @@ func (h *Handler) stopForward(local, remote string, port int) error {
} else { } else {
_ = h.Store.SetLinkDisabled(local, remote, port, true) _ = h.Store.SetLinkDisabled(local, remote, port, true)
} }
// The revoke task travels to whichever node OWNS the forward, and that node
// re-reads the disabled flag from its own store before starting a worker —
// so the flag has to be set on every node that has a copy of this link, not
// just the one handling this request. Propagating Disabled on the task lets
// the owner's RevokeFn stop the worker even if its own store row is stale.
//
// This also fixes a latent inconsistency: `ln` was read BEFORE the
// SetLinkDisabled(true) above, so the link published into the token still
// carried disabled=false and got copied into the topology entry verbatim.
ln.Disabled = true
// Publish the stop into the topology BEFORE revoking: the revoke retires the
// own-side worker/entry, so the flag must already exist somewhere that
// survives it and travels the ring. This is the send half of cluster-wide
// disabled propagation (see store.ReconcileLinkDisabled for the receive
// half).
if h.Ring != nil {
h.Ring.UpdateTopologyDisabled(local, remote, port, true)
}
if loc.LocalOnly { if loc.LocalOnly {
if h.Process != nil { if h.Process != nil {
key := process.WorkerKey(local, remote, port) key := process.WorkerKey(local, remote, port)

288
internal/store/link_test.go Normal file
View File

@ -0,0 +1,288 @@
package store
import (
"path/filepath"
"testing"
)
// seed inserts the local + remote rows a link's foreign keys require.
// (links.local / links.remote reference their own tables, so a link cannot
// exist on its own — the same reason ClaimFn upserts them before ReplaceLinks.)
func seed(t *testing.T, st *Store, locals []string, remote string) {
t.Helper()
for _, n := range locals {
if err := st.UpsertLocal(Local{Name: n, IP: "127.0.0.1", Port: 8080, Protocol: "tcp"}); err != nil {
t.Fatal(err)
}
}
if err := st.UpsertRemote(Remote{Name: remote, IP: "1.2.3.4", Port: 7000, Token: "tok", Enabled: true}); err != nil {
t.Fatal(err)
}
}
// TestLinkByTripleSurvivesReplaceLinks pins the reason LinkByTriple exists.
//
// ReplaceLinks() rewrites the whole table with DELETE + re-INSERT, so sqlite
// hands every row a FRESH autoincrement id. A Link captured before such a
// write (e.g. one riding inside a ring token) therefore carries an id that
// either matches a different forward or matches nothing. The natural key
// (local, remote, remotePort) is what every caller actually identifies a
// forward by, and it must survive those rewrites.
func TestLinkByTripleSurvivesReplaceLinks(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"alpha", "beta", "gamma"}, "srv")
links := []Link{
{Local: "alpha", Remote: "srv", RemotePort: 100},
{Local: "beta", Remote: "srv", RemotePort: 200},
{Local: "gamma", Remote: "srv", RemotePort: 300},
}
if err := st.ReplaceLinks(links); err != nil {
t.Fatal(err)
}
// Capture the ids as the ring would have them.
before := map[string]int64{}
all, err := st.ListLinks()
if err != nil {
t.Fatal(err)
}
for _, l := range all {
before[l.Local] = l.ID
}
if len(before) != 3 {
t.Fatalf("expected 3 links, got %d", len(before))
}
// Rewrite the table (this is what saveCanvas and ClaimFn both do).
if err := st.ReplaceLinks(links); err != nil {
t.Fatal(err)
}
after, err := st.ListLinks()
if err != nil {
t.Fatal(err)
}
if len(after) != 3 {
t.Fatalf("expected 3 links after rewrite, got %d", len(after))
}
// The natural key must still resolve to the right forward, with its
// disabled flag and group intact.
for _, l := range after {
if l.Disabled {
t.Errorf("link %s unexpectedly disabled after a plain rewrite", l.Local)
}
}
got, found, err := st.LinkByTriple("beta", "srv", 200)
if err != nil {
t.Fatal(err)
}
if !found {
t.Fatal("LinkByTriple failed to find beta after ReplaceLinks")
}
if got.Local != "beta" || got.RemotePort != 200 {
t.Fatalf("LinkByTriple returned the wrong row: %+v", got)
}
}
// TestLinkByTripleNotFoundIsNotError documents the contract callers rely on:
// "no persisted opinion yet" is (Link{}, false, nil), not an error. A fresh
// claim of a link with no row must be allowed to start.
func TestLinkByTripleNotFoundIsNotError(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
got, found, err := st.LinkByTriple("nope", "srv", 1234)
if err != nil {
t.Fatalf("missing link must not be an error, got %v", err)
}
if found {
t.Fatalf("missing link reported as found: %+v", got)
}
if got.Local != "" || got.RemotePort != 0 {
t.Fatalf("expected zero Link on miss, got %+v", got)
}
}
// TestLinkByTripleReadsDisabledFlag is the store-level half of the
// "stopped forwards resurrect on restart" bug: the claim path asks the store
// whether the user disabled this forward, so this lookup must return the flag
// as persisted.
func TestLinkByTripleReadsDisabledFlag(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"mc"}, "srv")
if err := st.ReplaceLinks([]Link{{Local: "mc", Remote: "srv", RemotePort: 25565, Group: "game"}}); err != nil {
t.Fatal(err)
}
if err := st.SetLinkDisabled("mc", "srv", 25565, true); err != nil {
t.Fatal(err)
}
ln, found, err := st.LinkByTriple("mc", "srv", 25565)
if err != nil {
t.Fatal(err)
}
if !found {
t.Fatal("expected to find the link")
}
if !ln.Disabled {
t.Fatal("expected Disabled=true to be visible through LinkByTriple")
}
if ln.Group != "game" {
t.Fatalf("group should survive, got %q", ln.Group)
}
// ...and the user-facing start path must be able to clear it again.
if err := st.SetLinkDisabled("mc", "srv", 25565, false); err != nil {
t.Fatal(err)
}
if ln, _, _ := st.LinkByTriple("mc", "srv", 25565); ln.Disabled {
t.Fatal("expected Disabled=false after clearing")
}
}
// TestGetLinkByIDIsStaleAfterReplaceLinks documents WHY callers must not use
// GetLink(id) with a previously captured id. It is not a fix — it is the trap
// being pinned shut, so the hazard stays visible if someone reintroduces it.
func TestGetLinkByIDIsStaleAfterReplaceLinks(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"alpha", "beta", "gamma"}, "srv")
if err := st.ReplaceLinks([]Link{
{Local: "alpha", Remote: "srv", RemotePort: 100},
{Local: "beta", Remote: "srv", RemotePort: 200},
}); err != nil {
t.Fatal(err)
}
all, _ := st.ListLinks()
var staleID int64
for _, l := range all {
if l.Local == "alpha" {
staleID = l.ID
}
}
if err := st.ReplaceLinks([]Link{
{Local: "alpha", Remote: "srv", RemotePort: 100},
{Local: "beta", Remote: "srv", RemotePort: 200},
{Local: "gamma", Remote: "srv", RemotePort: 300},
}); err != nil {
t.Fatal(err)
}
// The old id may still resolve, but to whatever row now occupies that
// id — which is exactly the silent-mis-target hazard. Assert that the
// natural key remains the only safe handle.
if ln, ok := st.GetLink(staleID); ok && ln.Local != "alpha" {
t.Logf("stale id %d now points at %q (hazard confirmed; use LinkByTriple)", staleID, ln.Local)
}
if got, found, _ := st.LinkByTriple("alpha", "srv", 100); !found || got.Local != "alpha" {
t.Fatalf("natural key must stay reliable, got %+v found=%v", got, found)
}
}
// TestReconcileLinkDisabledLearnsPeerDecision is the store half of
// cluster-wide stop propagation. A node that did NOT serve the stop request has
// no reason to know about it, and its links table is node-local — so the flag
// arrives via the ring and lands here. The case that matters is the OWNER of a
// forward on a different machine: before this existed, that node's copy still
// read "enabled", so it kept (or re-spawned) the worker for a forward the user
// had explicitly stopped.
func TestReconcileLinkDisabledLearnsPeerDecision(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"mc"}, "srv")
if err := st.ReplaceLinks([]Link{{Local: "mc", Remote: "srv", RemotePort: 25565, Group: "game"}}); err != nil {
t.Fatal(err)
}
// A peer's stop arrives.
if err := st.ReconcileLinkDisabled("mc", "srv", 25565, true); err != nil {
t.Fatal(err)
}
ln, found, err := st.LinkByTriple("mc", "srv", 25565)
if err != nil {
t.Fatal(err)
}
if !found {
t.Fatal("link disappeared during reconcile")
}
if !ln.Disabled {
t.Fatal("the peer's stop did not land in the local store")
}
if ln.Group != "game" {
t.Fatalf("reconcile must not clobber other fields, group=%q", ln.Group)
}
// A peer's re-enable arrives.
if err := st.ReconcileLinkDisabled("mc", "srv", 25565, false); err != nil {
t.Fatal(err)
}
if ln, _, _ := st.LinkByTriple("mc", "srv", 25565); ln.Disabled {
t.Fatal("the peer's re-enable did not land")
}
}
// TestReconcileLinkDisabledCreatesPlaceholderForUnknownForward: a node that has
// never seen the forward still must remember that it is stopped, otherwise a
// later claim on that node would resurrect it.
func TestReconcileLinkDisabledCreatesPlaceholderForUnknownForward(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"ghost"}, "srv")
if err := st.ReconcileLinkDisabled("ghost", "srv", 9999, true); err != nil {
t.Fatal(err)
}
ln, found, err := st.LinkByTriple("ghost", "srv", 9999)
if err != nil {
t.Fatal(err)
}
if !found {
t.Fatal("a stopped-but-unknown forward must be remembered, or a later claim resurrects it")
}
if !ln.Disabled {
t.Fatal("placeholder is not marked disabled")
}
}
// TestReconcileLinkDisabledIgnoresEnableForUnknown: an enable for a forward this
// node has never seen must NOT create a row. Creating one would invent forwards
// out of ring state.
func TestReconcileLinkDisabledIgnoresEnableForUnknown(t *testing.T) {
st, err := New(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatal(err)
}
defer st.Close()
seed(t, st, []string{"other"}, "srv")
if err := st.ReconcileLinkDisabled("unknown", "srv", 1234, false); err != nil {
t.Fatal(err)
}
if _, found, _ := st.LinkByTriple("unknown", "srv", 1234); found {
t.Fatal("an enable for an unknown forward must not materialise a row")
}
}

View File

@ -560,6 +560,30 @@ func (s *Store) GetLink(id int64) (Link, bool) {
return l, true return l, true
} }
// LinkByTriple looks a link up by its natural key (local, remote, remotePort).
//
// Prefer this over GetLink(id) whenever the caller only knows the forward's
// identity: ReplaceLinks() rewrites the whole table with DELETE + re-INSERT, so
// every row gets a fresh autoincrement id. Any id captured before such a write
// (e.g. a Link carried inside a ring token) is stale by definition and will
// either miss or — worse — match a different forward. The natural key is
// stable across those rewrites.
//
// Returns (link, found). A missing row is (Link{}, false) and is NOT an error:
// callers use that to mean "no persisted opinion yet".
func (s *Store) LinkByTriple(local, remote string, port int) (Link, bool, error) {
var l Link
err := s.db.QueryRow("SELECT id, local, remote, remote_port, offset_x, offset_y, grp, disabled FROM links WHERE local = ? AND remote = ? AND remote_port = ?", local, remote, port).
Scan(&l.ID, &l.Local, &l.Remote, &l.RemotePort, &l.OffsetX, &l.OffsetY, &l.Group, &l.Disabled)
if err != nil {
if errors.Is(err, sql.ErrNoRows) {
return Link{}, false, nil
}
return Link{}, false, err
}
return l, true, nil
}
// DeleteLink removes a single link by id. // DeleteLink removes a single link by id.
func (s *Store) DeleteLink(id int64) error { func (s *Store) DeleteLink(id int64) error {
_, err := s.db.Exec("DELETE FROM links WHERE id = ?", id) _, err := s.db.Exec("DELETE FROM links WHERE id = ?", id)
@ -639,6 +663,11 @@ func (s *Store) ReplaceLinks(links []Link) error {
// SetLinkDisabled flips the disabled flag of a forward identified by its // SetLinkDisabled flips the disabled flag of a forward identified by its
// (local, remote, remotePort) natural key. This is the persistence half of the // (local, remote, remotePort) natural key. This is the persistence half of the
// forwards-page start/stop toggle; the caller also drives the worker/ring side. // forwards-page start/stop toggle; the caller also drives the worker/ring side.
//
// Kept as a targeted UPDATE rather than a ReplaceLinks rewrite on purpose:
// ReplaceLinks deletes and reinserts every row, handing out fresh autoincrement
// ids and invalidating any Link a caller captured earlier (they travel inside
// ring tokens). Flipping one flag must not perturb other rows' identity.
func (s *Store) SetLinkDisabled(local, remote string, port int, disabled bool) error { func (s *Store) SetLinkDisabled(local, remote string, port int, disabled bool) error {
_, err := s.db.Exec( _, err := s.db.Exec(
"UPDATE links SET disabled = ? WHERE local = ? AND remote = ? AND remote_port = ?", "UPDATE links SET disabled = ? WHERE local = ? AND remote = ? AND remote_port = ?",
@ -647,6 +676,52 @@ func (s *Store) SetLinkDisabled(local, remote string, port int, disabled bool) e
return err return err
} }
// ReconcileLinkDisabled applies a cluster-wide view of one forward's disabled
// flag into the local store, creating a placeholder row when this node has none
// yet.
//
// This is the receive half of disabled-flag propagation. The stop decision is
// made on whichever node served the request, then rides the token ring in the
// topology entry; every other member calls this on adoption so its own
// links table agrees. Without it the flag lived only on the node that handled
// the request, and the node actually OWNS the forward — usually a different
// machine — still believed the forward was enabled and re-spawned its worker.
//
// A placeholder row is deliberate: a node that has never seen the forward still
// needs to remember "this is stopped" so a later claim on this node cannot
// resurrect it. The placeholder carries the same natural key, so a subsequent
// real claim fills in the rest.
func (s *Store) ReconcileLinkDisabled(local, remote string, port int, disabled bool) error {
cur, found, err := s.LinkByTriple(local, remote, port)
if err != nil {
return err
}
if found {
if cur.Disabled == disabled {
return nil // already agrees; avoid needless writes every token cycle
}
return s.SetLinkDisabled(local, remote, port, disabled)
}
if !disabled {
// Nothing to remember: an unknown forward with no entry is simply
// "not stopped", which is the default the claim path already assumes.
return nil
}
// Need a placeholder, which requires the local/remote foreign keys to exist.
if _, ok := s.GetLocal(local); !ok {
return nil // cannot materialise a link without its local peer row
}
if _, ok := s.GetRemote(remote); !ok {
return nil
}
links, err := s.ListLinks()
if err != nil {
return err
}
links = append(links, Link{Local: local, Remote: remote, RemotePort: port, Disabled: true})
return s.ReplaceLinks(links)
}
// SetLinkGroup assigns a management group label to a forward identified by its // SetLinkGroup assigns a management group label to a forward identified by its
// (local, remote, remotePort) natural key. Empty string clears the group // (local, remote, remotePort) natural key. Empty string clears the group
// (moves the forward to 未分组). This is the persistence half of the // (moves the forward to 未分组). This is the persistence half of the

View File

@ -2,11 +2,11 @@
// //
// Value is overridable at build time so release artifacts carry the real tag: // Value is overridable at build time so release artifacts carry the real tag:
// //
// go build -ldflags="-X webui4frpc/internal/version.Version=0.1.1" ./cmd/webui4frpc // go build -ldflags="-X webui4frpc/internal/version.Version=0.1.2" ./cmd/webui4frpc
// //
// Kept in its own package because both cmd/ and internal/httpapi report it // Kept in its own package because both cmd/ and internal/httpapi report it
// (main can't be imported, and httpapi must not depend on cmd). // (main can't be imported, and httpapi must not depend on cmd).
package version package version
// Version is the running build's semantic version, without a leading "v". // Version is the running build's semantic version, without a leading "v".
var Version = "0.1.1" var Version = "0.1.2"

View File

@ -1,5 +1,5 @@
Name: webui4frpc Name: webui4frpc
Version: %{?_w4f_version}%{!?_w4f_version:0.1.1} Version: %{?_w4f_version}%{!?_w4f_version:0.1.2}
Release: 1%{?dist} Release: 1%{?dist}
Summary: Visual frpc controller with token-ring cluster Summary: Visual frpc controller with token-ring cluster
License: MIT License: MIT
@ -67,6 +67,11 @@ fi
%doc %{_docdir}/webui4frpc/README.md %doc %{_docdir}/webui4frpc/README.md
%changelog %changelog
* Wed Sep 03 2026 JianFeeeee <jianfeeeee@gitcode> - 0.1.2-1
- 修复窄屏底部导航栏错位到顶部:顶栏的 backdrop-filter 会为
position:fixed 后代创建包含块,使 tab bar 的 bottom:0 相对顶栏而
非视口定位。窄屏下关掉顶栏模糊,改用近实色底。
* Wed Sep 02 2026 JianFeeeee <jianfeeeee@gitcode> - 0.1.1-1 * Wed Sep 02 2026 JianFeeeee <jianfeeeee@gitcode> - 0.1.1-1
- 修复令牌环冻结:心跳到达时复活被误判 offline 的前驱; - 修复令牌环冻结:心跳到达时复活被误判 offline 的前驱;
全环无可达后继时保留 inflight,让 leader 超时重发令牌。 全环无可达后继时保留 inflight,让 leader 超时重发令牌。

View File

@ -12,7 +12,7 @@
# 在 release.sh 流程内被调用;依赖 dist/artifacts/<version>/bin/<os>-<arch>/ 下的二进制。 # 在 release.sh 流程内被调用;依赖 dist/artifacts/<version>/bin/<os>-<arch>/ 下的二进制。
set -euo pipefail set -euo pipefail
VERSION="${1:-0.1.1}" VERSION="${1:-0.1.2}"
VERSION="${VERSION#v}" VERSION="${VERSION#v}"
ROOT="$(cd "$(dirname "$0")/.." && pwd)" ROOT="$(cd "$(dirname "$0")/.." && pwd)"
OUT="$ROOT/dist/artifacts/$VERSION" OUT="$ROOT/dist/artifacts/$VERSION"

View File

@ -3,7 +3,7 @@
# webui4frpc 多平台发布脚本 # webui4frpc 多平台发布脚本
# #
# 用法: # 用法:
# ./scripts/release.sh [version] # 默认 version=0.1.1 # ./scripts/release.sh [version] # 默认 version=0.1.2
# #
# 行为: # 行为:
# 1. 构建前端 (web/dist) 并同步到 internal/httpapi/dist (go:embed 用) # 1. 构建前端 (web/dist) 并同步到 internal/httpapi/dist (go:embed 用)
@ -19,7 +19,7 @@
# NO_BUILD_FRONTEND 设为 1 跳过前端构建 (用现有 dist) # NO_BUILD_FRONTEND 设为 1 跳过前端构建 (用现有 dist)
set -euo pipefail set -euo pipefail
VERSION="${1:-0.1.1}" VERSION="${1:-0.1.2}"
VERSION="${VERSION#v}" # 去掉可能的前导 v VERSION="${VERSION#v}" # 去掉可能的前导 v
ROOT="$(cd "$(dirname "$0")/.." && pwd)" ROOT="$(cd "$(dirname "$0")/.." && pwd)"
OUT="$ROOT/dist/artifacts/$VERSION" OUT="$ROOT/dist/artifacts/$VERSION"

View File

@ -2,7 +2,8 @@
"""上传 release 资产到 gitcode(两步:取签名 URL → PUT 到 OBS)。 """上传 release 资产到 gitcode(两步:取签名 URL → PUT 到 OBS)。
用法: upload_assets.py <version> <token> [file...] 用法: upload_assets.py <version> <token> [file...]
不传 file 时上传该版本目录下所有 tar.gz/zip 与 SHA256SUMS。 不传 file 时上传该版本目录下全部发布产物:
tar.gz / zip / deb / rpm / pkg / -setup.exe / SHA256SUMS
""" """
import json import json
import os import os
@ -38,6 +39,19 @@ def put_file(url: str, headers: dict, path: str) -> tuple[int, str]:
return e.code, e.read().decode("utf-8", "replace")[:300] return e.code, e.read().decode("utf-8", "replace")[:300]
# 发布产物后缀:免安装包 + 原生安装包。
# 注意 -setup.exe 而不是裸 .exe,避开 bin/ 里的裸二进制。
ARTIFACT_SUFFIXES = (
".tar.gz", ".zip", # 免安装
".deb", ".rpm", ".pkg", # linux / macOS 安装包
"-setup.exe", # windows 安装器
)
def is_artifact(name: str) -> bool:
return name == "SHA256SUMS" or name.endswith(ARTIFACT_SUFFIXES)
def main() -> int: def main() -> int:
if len(sys.argv) < 3: if len(sys.argv) < 3:
print(__doc__) print(__doc__)
@ -47,10 +61,7 @@ def main() -> int:
os.path.dirname(os.path.dirname(os.path.abspath(__file__))), os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
"dist", "artifacts", version, "dist", "artifacts", version,
) )
files = sys.argv[3:] or sorted( files = sys.argv[3:] or sorted(f for f in os.listdir(outdir) if is_artifact(f))
f for f in os.listdir(outdir)
if f.endswith((".tar.gz", ".zip")) or f == "SHA256SUMS"
)
failed = 0 failed = 0
for name in files: for name in files:
path = os.path.join(outdir, name) path = os.path.join(outdir, name)

View File

@ -384,16 +384,37 @@ const currentLabel = computed(() => nav.value.find((n) => n.key === view.value)?
align-items: center; align-items: center;
gap: 10px; gap: 10px;
padding: 10px 12px; padding: 10px 12px;
overflow: visible; overflow: hidden;
/* 兵家必争:任何子项算错都不得弄出横向滚动条 */
overflow-x: hidden;
border-right: none; border-right: none;
border-bottom: 1px solid var(--w4f-line); border-bottom: 1px solid var(--w4f-line);
padding-top: max(10px, env(safe-area-inset-top)); padding-top: max(10px, env(safe-area-inset-top));
/* 关键:backdrop-filter 会为 position:fixed 后代创建【包含块】
(与 transform/filter 同理)。若保留它,下面 .sb-nav 的
bottom:0 就是相对这条顶栏定位而非视口,底部 tab bar 会贴在
顶栏下沿、看起来固定在屏幕顶部。因此窄屏必须关掉顶栏的玻璃
模糊,改用接近实色的底,模糊效果留在 .sb-nav 自己身上
(元素自身的 backdrop-filter 不影响它自己的定位)。 */
backdrop-filter: none;
-webkit-backdrop-filter: none;
background: color-mix(in srgb, var(--w4f-card-solid) 92%, transparent);
} }
.sb-brand { .sb-brand {
padding: 0; padding: 0;
gap: 8px; gap: 8px;
flex: 1 1 auto; flex: 1 1 auto;
min-width: 0; min-width: 0;
overflow: hidden;
}
/* 品牌文字必须可收缩:“webui4frpc”是不可断行单词,不给 min-width:0
+ overflow:hidden 的话它的 flex 自动最小尺寸等于 min-content(≨85px),
顶栏就拿不出空间给右侧身份区,溢出为横向滚动。 */
.sb-brand-text { min-width: 0; overflow: hidden; }
.sb-brand-text b {
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
} }
.sb-logo { width: 30px; height: 30px; border-radius: 10px; } .sb-logo { width: 30px; height: 30px; border-radius: 10px; }
.sb-brand-text b { font-size: 14.5px; } .sb-brand-text b { font-size: 14.5px; }
@ -425,7 +446,10 @@ const currentLabel = computed(() => nav.value.find((n) => n.key === view.value)?
gap: 2px; gap: 2px;
padding: 5px 2px; padding: 5px 2px;
border-radius: 11px; border-radius: 11px;
font-size: 10.5px; /* Keep the tab label at the theme's 11px floor; the 10.5px here (and 9.5px
* at ≤380px) read as noise on a phone. "账号与密钥" is still handled by the
* .label ellipsis (width:100% + overflow:hidden). */
font-size: 11px;
font-weight: 600; font-weight: 600;
text-align: center; text-align: center;
} }
@ -442,23 +466,30 @@ const currentLabel = computed(() => nav.value.find((n) => n.key === view.value)?
box-shadow: none; box-shadow: none;
} }
/* 身份与主题折进顶条右侧 */ /* 身份与主题折进顶条右侧。
不再用 vw 限宽:vw 含垂直滚动条宽度,会比实际可用内容宽大;而且
max-width 只封顶盒子,内部 flex:0 0 auto 的徒章/按钮仍会撑出去。
改为全链路可收缩(min-width:0 + overflow:hidden),由 flex 自行分配。 */
.sb-foot { .sb-foot {
margin-top: 0; margin-top: 0;
padding-top: 0; padding-top: 0;
display: flex; display: flex;
align-items: center; align-items: center;
gap: 8px; gap: 8px;
flex: 0 0 auto; flex: 0 1 auto;
min-width: 0;
} }
.sb-identity { .sb-identity {
margin-bottom: 0; margin-bottom: 0;
padding: 5px 8px; padding: 5px 8px;
font-size: 11.5px; font-size: 11.5px;
max-width: 42vw; flex: 0 1 auto;
min-width: 0;
max-width: none;
overflow: hidden;
} }
.sb-identity .id-name { max-width: 14vw; } .sb-identity .id-name { min-width: 0; max-width: none; }
.sb-theme { margin-bottom: 0; } .sb-theme { margin-bottom: 0; flex: 0 0 auto; }
.sb-theme .sw { width: 22px; height: 22px; border-radius: 7px; } .sb-theme .sw { width: 22px; height: 22px; border-radius: 7px; }
.main { height: auto; flex: 1; min-height: 0; } .main { height: auto; flex: 1; min-height: 0; }
@ -470,12 +501,15 @@ const currentLabel = computed(() => nav.value.find((n) => n.key === view.value)?
} }
} }
/* 极窄(≤ 380px,iPhone SE / 小屏安卓):进一步压缩顶条 */ /* 极窄(≤ 380px,iPhone SE / 小屏安卓):进一步压缩顶条。
角色徒章是 flex:0 0 auto 不可收缩的(“超级管理员”≨70px),极窄下隐去,
用户名与退出按钮保留。
注意:底部 tab 文字不再随之降到 9.5px —— 它在手机上低于可读下限
(见 .sb-i 的 11px 注释),长标签交给 .label 的省略号处理。 */
@media (max-width: 380px) { @media (max-width: 380px) {
.sb-brand-text b { font-size: 13px; } .sb-brand-text b { font-size: 13px; }
.sb-identity { max-width: 38vw; padding: 4px 7px; } .sb-identity { padding: 4px 7px; }
.sb-identity .id-level { display: none; } .sb-identity .id-level { display: none; }
.sb-i { font-size: 9.5px; }
.content { padding-left: 10px; padding-right: 10px; } .content { padding-left: 10px; padding-right: 10px; }
} }
</style> </style>

View File

@ -184,6 +184,20 @@
/* ---- blue theme (ocean × frost) ---- */ /* ---- blue theme (ocean × frost) ---- */
/* ---- box model reset ----
* This app had NO box-sizing reset, so every element using `width: 100%` with
* horizontal padding overflowed by exactly its padding under the default
* content-box model. The narrow-screen top bar was the visible victim:
* `.sidebar { width: 100%; padding: 10px 12px }` rendered 375+24 = 399px wide
* on a 375px viewport (344px at 320px), pushing the theme swatches and logout
* button off the right edge where `overflow-x: hidden` silently clipped them.
* It also made the desktop sidebar 246+28 = 274px wide instead of 246px.
* border-box makes width include padding+border, which every layout here
* already assumes (Element Plus sets it on its own components already). */
*,
*::before,
*::after { box-sizing: border-box; }
html, html,
body, body,
#app { #app {
@ -346,4 +360,14 @@ a { color: var(--w4f-primary-h); }
/* ---- Element Plus dialog tweaks to read as glass ---- */ /* ---- Element Plus dialog tweaks to read as glass ---- */
.el-overlay { backdrop-filter: blur(3px); } .el-overlay { backdrop-filter: blur(3px); }
.el-dialog { border-radius: var(--w4f-radius) !important; box-shadow: var(--w4f-sh-lg) !important; } .el-dialog { border-radius: var(--w4f-radius) !important; box-shadow: var(--w4f-sh-lg) !important; }
/* Dialogs pass a fixed px width (width="420px" / "440px"), which Element Plus
* writes into an INLINE `--el-dialog-width` custom property. On phones that is
* wider than the viewport (420px on a 375px screen), so the dialog's right edge
* — including the header close button — starts off-screen and is only reachable
* by scrolling the overlay horizontally. Clamp it to the viewport; !important
* is required to beat the inline custom property. */
@media (max-width: 720px) {
.el-dialog { --el-dialog-width: calc(100vw - 24px) !important; }
}
.el-message { border-radius: var(--w4f-radius-sm) !important; box-shadow: var(--w4f-sh-lg) !important; } .el-message { border-radius: var(--w4f-radius-sm) !important; box-shadow: var(--w4f-sh-lg) !important; }

View File

@ -25,16 +25,17 @@
<VueFlow <VueFlow
v-model:nodes="nodes" v-model:nodes="nodes"
v-model:edges="edges" v-model:edges="edges"
:default-viewport="{ zoom: 0.85 }" :default-viewport="{ zoom: FIT_MAX_ZOOM }"
:min-zoom="0.3" :min-zoom="MIN_ZOOM"
:max-zoom="2" :max-zoom="2"
fit-view-on-init
:delete-key-code="null" :delete-key-code="null"
:nodes-draggable="canWrite" :nodes-draggable="canWrite"
:nodes-connectable="canWrite" :nodes-connectable="canWrite"
:edges-updatable="false" :edges-updatable="false"
:edges-reconnectable="false" :edges-reconnectable="false"
class="flow-canvas" class="flow-canvas"
@init="onFlowInit"
@nodes-initialized="reframe"
@connect="onConnect" @connect="onConnect"
@pane-click="deselectAll" @pane-click="deselectAll"
@edge-click="onEdgeClick" @edge-click="onEdgeClick"
@ -118,7 +119,7 @@
<script setup lang="ts"> <script setup lang="ts">
import '@vue-flow/core/dist/style.css' import '@vue-flow/core/dist/style.css'
import '@vue-flow/core/dist/theme-default.css' import '@vue-flow/core/dist/theme-default.css'
import { ref, computed, onMounted, onBeforeUnmount } from 'vue' import { ref, computed, onMounted, onBeforeUnmount, nextTick } from 'vue'
import { VueFlow } from '@vue-flow/core' import { VueFlow } from '@vue-flow/core'
import { Background } from '@vue-flow/background' import { Background } from '@vue-flow/background'
import { ElMessage, ElMessageBox } from 'element-plus' import { ElMessage, ElMessageBox } from 'element-plus'
@ -134,6 +135,103 @@ const nodes = ref<any[]>([])
const edges = ref<Edge[]>([]) const edges = ref<Edge[]>([])
const loading = ref(false) const loading = ref(false)
// ---- viewport / zoom floor ----
// The canvas is laid out in absolute graph units (local column at x=60, remote
// column at x=760, and BOTH columns grow downward on every added forward), so
// its bounding box is far wider and taller than a phone viewport. Vue Flow's
// `fit-view-on-init` frames that whole box, which on a phone produced a scale of
// ~0.31: node cards measure 248px in graph units but rendered 85px wide at an
// effective font size of ~4.4px — unreadable. Even a 1280px desktop landed at
// ~0.60 (8.4px) with five forwards stacked.
//
// So the canvas is framed by frameCanvas() below, clamped by a zoom floor, and
// the canvas pans when the content no longer fits. 0.65 renders a 240px local
// card at ~156px with a legible ~9px font.
//
// NOTE: `fit-view-on-init` is deliberately NOT set. It resolves on its own
// schedule and overwrote the explicit fit performed here (observed: this code
// measured the unfitted layout, set the viewport, and then the deferred init fit
// landed last and won). Owning the fit outright is the only way to keep it
// deterministic.
//
// The store arrives via the `init` event. That is deliberately preferred over
// `useVueFlow()`: this component is the PARENT of <VueFlow>, and calling
// useVueFlow() in a parent without a shared id can resolve to a different store
// than the one the rendered flow uses. The event cannot misfire that way.
const MIN_ZOOM = 0.65
const FIT_MAX_ZOOM = 1
type FlowStore = {
fitView: (opts?: Record<string, unknown>) => Promise<boolean>
getViewport?: () => { x: number; y: number; zoom: number }
setViewport?: (
t: { x: number; y: number; zoom: number },
opts?: Record<string, unknown>,
) => Promise<boolean>
}
let flowStore: FlowStore | null = null
// frameCanvas fits the graph, then re-anchors it to the top-left padding edge.
//
// Two reasons it is not just fitView():
// 1. the zoom floor (MIN_ZOOM) — fitView alone shrank 5 stacked forwards to
// ~0.31 on a phone, rendering 4.4px text;
// 2. fitView centres the bounding box, which on a narrow screen puts the seam
// between the local and remote columns mid-viewport — the user sees half of
// each. Anchoring the top-left node to the padding edge means whole local
// cards are visible immediately and the remote column is a pan away.
//
// duration:0 on both transitions — we measure node rects in between, and a
// running animation would hand back mid-flight transforms and skew the anchor.
const PANE_PAD_RATIO = 0.06
const frameCanvas = async () => {
if (!flowStore) return
await flowStore.fitView({
padding: 0.12,
minZoom: MIN_ZOOM,
maxZoom: FIT_MAX_ZOOM,
duration: 0,
})
const pane = document.querySelector<HTMLElement>('.vue-flow')
const nodeEls = Array.from(document.querySelectorAll<HTMLElement>('.vue-flow__node'))
if (!pane || !nodeEls.length) return
const paneRect = pane.getBoundingClientRect()
const rects = nodeEls.map((el) => el.getBoundingClientRect())
const dx =
paneRect.left + pane.clientWidth * PANE_PAD_RATIO - Math.min(...rects.map((r) => r.left))
const dy =
paneRect.top + pane.clientHeight * PANE_PAD_RATIO - Math.min(...rects.map((r) => r.top))
if (Math.abs(dx) < 1 && Math.abs(dy) < 1) return
const vp = flowStore.getViewport?.()
if (!vp) return
await flowStore.setViewport?.({ x: vp.x + dx, y: vp.y + dy, zoom: vp.zoom }, { duration: 0 })
}
// The flow store and the node list become available at different times: the
// store at component init, the nodes only once the async canvas fetch resolves
// (plus a further tick for Vue Flow to measure them). Framing on any single
// signal would fit an EMPTY canvas, so every path calls reframe() and it no-ops
// until there is a store and something to frame.
//
// `nodes-initialized` also fires whenever a node is added later. Auto-framing on
// that would yank the viewport out from under whoever is mid-pan, so the initial
// frame is claimed exactly once; rearranging is the only thing that re-frames on
// demand (see autoLayout).
let framedOnce = false
const reframe = (force = false) => {
if (!flowStore || !nodes.value.length) return
if (framedOnce && !force) return
framedOnce = true
nextTick(() => { void frameCanvas() })
}
const onFlowInit = (store: FlowStore) => {
flowStore = store
reframe()
}
// Background dot-pattern color, resolved from the live theme var so the grid // Background dot-pattern color, resolved from the live theme var so the grid
// recolors on theme switch (vue-flow Background renders pattern-color as an // recolors on theme switch (vue-flow Background renders pattern-color as an
// SVG fill attribute, which can't resolve var() directly — read the computed // SVG fill attribute, which can't resolve var() directly — read the computed
@ -667,7 +765,12 @@ const remoteNodes = computed(() =>
nodes.value.filter((n) => n.type === 'remote'), nodes.value.filter((n) => n.type === 'remote'),
) )
const autoLayout = () => layoutColumns(true) // Re-frame after rearranging, bounded by the same zoom floor so "自动排列" never
// shrinks the cards back into unreadable territory on a narrow screen.
const autoLayout = () => {
layoutColumns(true)
reframe(true)
}
const load = async () => { const load = async () => {
loading.value = true loading.value = true
@ -718,6 +821,9 @@ const load = async () => {
} }
assignLayers() assignLayers()
dirty.value = false dirty.value = false
// The store exists by now but the node list has only just been populated, so
// this is the first moment a fit actually has something to frame.
reframe(true)
} finally { } finally {
loading.value = false loading.value = false
} }

View File

@ -481,7 +481,11 @@ onBeforeUnmount(() => {
</script> </script>
<style scoped lang="scss"> <style scoped lang="scss">
.cluster-page { height: 100%; display: flex; flex-direction: column; overflow-y: auto; padding: 4px 0 40px; gap: 14px; } /* Let .content (App.vue) remain the single scroll container — see the note in
* StatusView: a page-level scroller sized by leftover flex space is what froze
* the status page. `.log-list` keeps its own bounded scroll (that one is
* deliberate — a log pane should not push the whole page). */
.cluster-page { display: flex; flex-direction: column; padding: 4px 0 40px; gap: 14px; }
/* hero */ /* hero */
.hero { display: flex; align-items: center; justify-content: space-between; gap: 16px; flex-wrap: wrap; padding: 16px 20px; } .hero { display: flex; align-items: center; justify-content: space-between; gap: 16px; flex-wrap: wrap; padding: 16px 20px; }

View File

@ -167,11 +167,12 @@ onMounted(load)
</script> </script>
<style scoped lang="scss"> <style scoped lang="scss">
/* Let .content (App.vue) remain the single scroll container — see the note in
* StatusView: a page-level scroller sized by leftover flex space can collapse
* to zero height on narrow screens. */
.settings-page { .settings-page {
height: 100%;
display: flex; display: flex;
flex-direction: column; flex-direction: column;
overflow-y: auto;
} }
.page-top { .page-top {

View File

@ -496,11 +496,22 @@ onBeforeUnmount(() => {
</script> </script>
<style scoped lang="scss"> <style scoped lang="scss">
/* The pages inside .content must NOT be their own scrollers.
*
* App.vue already designates `.content` as the app's single scroll container
* (`flex:1; overflow-y:auto`). These pages were also `height:100%` +
* `overflow-y:auto`, producing a nested scroller whose height is the leftover
* space of a flex column. On a 375px phone the status page's fixed-height
* header (122px) plus the single-column KPI stack (565px) exceed the available
* 651px, so the inner `.status-body` was handed 0px while holding 1682px of
* content — the page looked completely frozen, since neither `.content` (which
* had nothing to scroll) nor the 0-height inner box could be swiped.
*
* Letting the page grow naturally and scroll in `.content` removes the whole
* class of bug: there is no leftover-space arithmetic left to get wrong. */
.status-page { .status-page {
height: 100%;
display: flex; display: flex;
flex-direction: column; flex-direction: column;
overflow: hidden;
gap: 14px; gap: 14px;
} }
@ -527,12 +538,9 @@ onBeforeUnmount(() => {
} }
.status-body { .status-body {
flex: 1;
display: flex; display: flex;
flex-direction: column; flex-direction: column;
gap: 18px; gap: 18px;
overflow-y: auto;
padding-right: 4px;
} }
.section { display: flex; flex-direction: column; } .section { display: flex; flex-direction: column; }
@ -720,8 +728,13 @@ onBeforeUnmount(() => {
} }
@media (max-width: 380px) { @media (max-width: 380px) {
/* 极窄:KPI 也降为单列,否则数字会被压断 */ /* Keep the KPI grid at TWO columns even on the narrowest phones.
.kpis { grid-template-columns: minmax(0, 1fr); } * Dropping to one column made four full-width cards ~565px tall, which on its
* own exceeded the whole viewport and (with the page-level scroller of the
* day) collapsed the content area to zero height. Two columns keeps each card
* ~160px wide — enough for the 27px number + 11.5px caption, which is all a
* KPI card holds. */
.kpis { gap: 8px; }
.fwd-route { font-size: 12.5px; } .fwd-route { font-size: 12.5px; }
.page-title { font-size: 16px; } .page-title { font-size: 16px; }
} }