Commit Graph

79 Commits

Author SHA1 Message Date
cf14f66ad4 fix(gateway): AUTO 被密钥模型范围拦下时报错误导 + 编辑器给出提示
一个密钥的模型范围不含 AUTO 时,它用不了 AUTO —— 范围过滤发生在链之前。
但这两条路径都不说真话:

1. 请求侧:AUTO 走的是通用分支,返回 model_not_found "model \"AUTO\" is
   not configured",读起来像 AUTO 没配置。非 AUTO 模型早就有 model_not_allowed
   的专门提示,AUTO 漏了。补上,并说明修法(把 AUTO 加进该密钥的 models,
   或去掉范围限制)。

2. 编辑器侧:per-key AUTO 链编辑器可以正常配链、保存也成功,看起来一切正常,
   但该密钥的每个请求都会 403。管理员无从得知。现在弹窗顶部在检测到冲突时
   显示警告,并列出当前范围。

刻意不做的事:不自动把 AUTO 加进该密钥的模型范围。那等于悄悄授予运营
没要求的访问权,比一个显眼的警告更糟。

判据 1 条,除确认提示出现外还断言该函数体内没有 PUT/POST/fetch/api ——
它只报告,不得写回。变异(去掉 AUTO 判断)判红。

CDP 实测四种场景:范围含 AUTO → 无提示;范围不含 → 警告并列出范围;
空范围(不受限)→ 无提示;范围含 AUTO → 无提示。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 21:23:02 +08:00
130b3f2e8d feat(ui): per-key AUTO 链默认复制全局链作为编辑起点
此前给没有独立链的密钥打开编辑器是空白画布,管理员要从 226 个模型里
一条条重新挑。真实任务是"这个用户拿全局链,但去掉他不该碰的两个源",
所以起点应该是全局链本身。

- 无独立链时读 /api/auto 的 rules 预填(实测 4 泳道 / 10 块,与全局链
  一致),不是 discovery 兜底(那会把每个源的每个模型都堆进来)。
- 提示区分三种状态,因为它们在画布上长得一样、点取消的后果完全不同:
  复制来的起点 / 该密钥已有的独立链 / 清空后保存将恢复跟随全局。
- seeded 标记随关闭清理,避免下次的提示说错来源。

CDP 实测:打开无独立链的密钥 → 画布 4 泳道 10 块 + 提示"已复制全局链
作为编辑起点";删一个块服务端仍 inherits:true;保存后 inherits:false
n=9、弹窗关闭;重开读回 4 泳道 9 块且提示消失(它现在有自己的链了)。
横向滚动可达性一并验证:scrollLeft 172 = scrollWidth-clientWidth,
滚到底后无残留不可达块。

判据 2 条 + 变异(恢复空白起点 / 去掉 seeded 标记均判红)。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 13:36:11 +08:00
dac0b4ce93 fix(ui): 统计区间文案在日视图恒为空、天数按小时桶误报
v1.9.0 加的「统计区间:起 → 止 (Nd)」正是用来解释"周 > 月"的,但它自己
在日视图上从来不出现在屏幕上。

原因:区间终点从最后一个 bucket 的标签拼出来,而 bucket 粒度随视图变化
—— 日视图按小时("2026-10-03T09"),周/月按天("2026-10-01")。代码无条件
拼 "T00:00:00Z",日视图就成了 "2026-10-03T09T00:00:00Z",Date 拒绝该值,
toISOString() 抛 RangeError,整个赋值语句丢失。页面无任何报错,只是那行
字不见了——恰好是用户最常看的"今日"视图。

第二个问题同源:天数直接用了 st.buckets.length。日视图 5 个桶是 5 个小时,
界面却显示 "(5d)"。改为按窗口起止算天数。

CDP 实测(修复后):
  day   统计区间: 2026-10-03 → 2026-10-03 (1d)   ← 原为空,且曾误报 5d
  week  统计区间: 2026-09-28 → 2026-10-03 (6d)
  month 统计区间: 2026-10-01 → 2026-10-03 (3d)
  all  (空,正确:终身累计没有窗口)

判据一条 + 变异:恢复原来的盲目拼接 → 判红。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 13:36:11 +08:00
f9cd0770a9 fix(ui): per-key 弹窗复用共享弹窗ID、编辑即落盘、空画布显示226个模型
截图看渲染 + 交互实测后发现的四个问题,都出在我自己新写的代码里。

一、弹窗用了全站共享的 #modal-wrap。
这个 id 是 App 内所有弹窗共用的,且允许多个同时存在("添加档位"选择器就
开在本弹窗之上)。既有代码靠 .closest() 或 closeTopModal() 从按钮往上找,
是有效的——但本弹窗还要从画布侧被查(sortCanvasEl / keyAutoClose /
keyAutoInherit),那里没有按钮可走。实测点"添加档位"后 DOM 里同时有
modal-wrap=1 和 key-auto-modal=1,getElementById 会命中文档里的第一个,
也就是用户没在看的那一个:取消关掉错误的框,画布也可能解析到选择器上。
改为独立 id #key-auto-modal。

二、编辑即落盘,保存/取消形同虚设。
scrAddFromForm / sortScopeSave / scrDelSlot 等 6 处编辑后立刻调
persistAuto()。全局页这样是有意的(改动立即生效),但它和自己页面上那句
"点击保存排序后生效"的提示自相矛盾;在 per-key 弹窗里更严重——加一个槽位
就已经写进那个用户的链了,取消也无法撤销。收敛到 afterChainEdit():
有 scope 时只提示、不落盘,保存是唯一写入口。

三、无独立链时画布显示全部 226 个模型。
注释写的是"显示空画布",实现却是 buildLanes(idx, [], 'chat')——空 rules
会走 discovery 兜底,把每个源的模型全填进去。用户打开弹窗看到 226 个色块,
会以为这些就是这个密钥的链。改为空 rules 直接给 []。

四、"使用全局链"提示不随画布变化。
加完槽位后仍显示"该密钥使用全局 AUTO 链",与旁边的色块直接矛盾。改为由
afterChainEdit 触发、按画布是否有内容决定去留。

另:批量替换时正则的缩进假设错了一次,把 afterChainEdit 自己改成了自调用
(无限递归),node --check 抓到后修掉。这类批量替换必须先过语法检查。

判据两条 + 变异:弹窗回到共享 id → 判红;去掉 scope 分支 → 判红。
CDP 实测:弹窗 id 唯一、添加后画布 0→1 块、添加后服务端 inherits=true
(未落盘)、取消后仍 inherits=true、选择器与弹窗并存互不干扰。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 11:20:29 +08:00
6311913812 fix(stats): 日/周/月视图请求记录表空白,滚动加载顺序错乱
首页记录表由 paintRecords(st.records) 渲染,而 /api/stats 的 period 分支
手写了一份响应 map,压根没有 records 字段 —— 只有 all 分支走 Snapshot()
才带得上。实测切到日/周/月,记录表恒为空。

两处数据源也不一致:

| 端点 | 数据源 | 排序 |
|---|---|---|
| /api/stats(all)| 内存ring(几百条)| oldest-first |
| /api/stats/records(翻页)| 审计文件(全部)| newest-first |
| /api/stats(period)| 无 | 无 |

前端靠 paintRecords 里一句 .reverse() 把 all 的 oldest-first 转成 newest-first,
翻页端点则本来就是 newest-first。首屏和翻页方向相反,滚动时新旧行会错位
拼接;且 all 分支不带 next_cursor,表永远停在第一屏。

修法:两个分支都改用 AuditPage(与翻页端点同一份代码路径),记录方向统一
由它决定,前端去掉 .reverse()。附带好处:记录深度不再受 maxRecs 限制,
"全部"视图真的能翻到底。

判据三条,变异验证:period 分支返回 nil records → 前两条判红;all 分支不
覆盖 records → "no next_cursor"判红。变异过程中我自己也犯了两次错:第一版
判据硬编码 newest-first,而AuditPage 在单文件(整块读完)与多文件(分块反向
读)下方向不同,改为断言两处方向一致;第三版想测 admin 过滤,但
newTestGateway 的 sk-test 不是 admin,改为测真正的越权边界(非admin 不能
用 ?key= 看别人的记录)。

CDP 实测:四个视图各 100 行、时间从新到旧,滚动加载 100→200 行且新末行
时间更旧(顺序衔接正确)。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 10:32:51 +08:00
aa003e84d3 fix(ui): per-key 编辑器画布画错容器、保存成功却不关弹窗
CDP 实测发现两个真 bug,都是复用画布时引入的。

一、画布画进了后台标签页。sortCanvasEl() 原本按文档顺序找,先命中
#scr-canvas(全局优先级页的画布)。而弹窗是挂在设置页之上的,两个画布
同时存在于 DOM —— 于是 per-key 的 36 条泳道被画进隐藏的全局画布,弹窗
里空空如也,回���全局页还会发现自己的链被覆盖了。

改为按 sortState.scope 解析:有 scope 就用 #key-auto-canvas,否则用
#scr-canvas。

二、保存成功却弹窗不关。!j.slots 分支里我写了 return,跳过了
keyAutoClose。写确实成功了(服务端已存),但用户看到弹窗还开着、按钮
还是禁用状态,会反复点保存。而且那条 toast 文案写成"保存后 AUTO 请求
会失败",读起来像保存失败了 —— 一并改成"已保存,但槽位无法解析…"。

判据两条,变异验证:恢复按文档顺序查找 → 第一条判红;!slots 分支加
return → 第二条判红。第二条最初只查 keyAutoClose() 是否存在,变异加了
return 仍然通过,属于判据漏放,已改为扫描 try 块内 close 调用之前的
所有 return 语句。

CDP 复测(先开全局页制造干扰,再开 per-key 弹窗):
- 弹窗画布 36 泳道,全局画布保持 4 泳道未被污染
- 保存 gpt-5.2@xinjianya 后弹窗关闭、服务端存储正确、scope 清空
- 重开弹窗读回 [["gpt-5.2@xinjianya"]]

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 10:32:51 +08:00
13e2397e9c fix(ui): per-key 链复用全局 AUTO 画布;统计周期显示实际窗口
两件事。

一、per-key AUTO 链的 UI 之前自己写了一套表格,只有 model + tier 两列,
比全局 AUTO 编辑器弱:没有拖拽排序、没有每槽位的 quota/period/hours。
这是重复实现且必然漂移。改为复用同一套泳道画布:

- 抽出 buildSortIndex / buildLanes / lanesToRules / renderSortEditor,
  全局编辑器与 per-key 编辑器共用。
- sortState.scope 决定保存目标(null=全局,key=该密钥),persistAuto
  按此分支;空 lane 列表即"恢复继承",与后端空 PUT 等价。
- 画布查找走 sortCanvasEl(),两个 id 都能解析,否则拖拽在弹窗里静默失效。
- keyAutoClose 必须清 scope,否则关掉弹窗后全局编辑器会存到这个 key。
- 删掉旧的 keyAutoRow/keyAutoAddRow/keyAutoClear。

判据改为断言"复用"而非断言具体实现,并新增两条:
- 关闭弹窗不清 scope(全局编辑器会存错目标)
- per-key 编辑器不再挂载共享画布

二、"周视图总量大于月视图"不是 bug:日历周期并非嵌套。"本周"从周一开始,
所以月初会回溯到上月末几天,本周总量合法地大于本月。实测今天 2026-10-03
(周六):周窗口 09-28→10-03(6 天,9.2B tokens),月窗口 10-01→10-03
(3 天,4.0B tokens)。

真正的缺陷是看不见这一点。统计页新增"统计区间:YYYY-MM-DD → YYYY-MM-DD
(Nd)",让周 > 月 自解释。判据钉住这个事实,并做变异验证:把周窗口钳制到
月内(一个看起来很自然的"修复")会被两条判据同时抓住。

CDP 实测:周视图显示"2026-09-28 → 2026-10-03 (6d)",月视图显示
"2026-10-01 → 2026-10-03 (3d)";per-key 编辑器渲染出 36 条泳道 / 226 个
块,与全局编辑器同款。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 08:26:57 +08:00
22decf2bd3 feat(keys): 每个密钥可配独立 AUTO 链,管理员代配
此前 AUTO 链是全局单值(cfg.Auto + Core.AutoChain()),所有用户共用一条链。
管理员无法为某个用户单独指定调度链。

按 per-key 覆盖 + 全局兜底实现:

- config.GWKey 增加 Auto 字段与 HasOwnAuto()。未配置即继承全局链,
  存量部署零改动,新密钥天然继承全局链。
- Core 把 buildAutoChain 的归一化逻辑抽成 chainForRules,全局链、
  生图链、per-key 链共用同一套编译,避免两处漂移。
- Core 增加 keyAutoChains 缓存 + AutoChainFor(key)。请求路径读缓存不
  加锁,与全局链查询一致。缓存整表原子替换,不会看到半成品。
- 请求侧 chat.go 改用 AutoChainFor(reqKey)。冷却仍在 Provider 上按
  model+source 共享:两条链指向同一个 slot 时共用冷却,与今天单链行为
  相同,也避免为 per-key 维度重构冷却而改变现有可观测语义。
- API:GET/PUT/DELETE /api/keys/{key}/auto(admin),GET
  /api/keys/me/auto(任意角色,只能读自己的)。空 PUT 与 DELETE 等价于
  "恢复继承",无法持久化一条会 503 的空链。写入时回报解析出的槽位数,
  让管理员当场看到模型名写错,而不是等用户下次请求 503。
- 源变更时一并重编译 per-key 链,加源后无需重启即可生效。

判据 18 条,5 个变异全部被抓住:AutoChainFor 忽略 key、清空后不重建
缓存、源变更不重建、空 PUT 落盘成空链、me/auto 误要求 admin。

UI 判据做变异时发现漏放:只查函数定义存在,删掉按钮后仍通过。已改为
断言 keyCanvasHtml 内的调用点。

端到端实测(真实 HTTP + 两个 mock 上游):admin 与 bob 初始同为 m-fast,
给 bob 配 m-cheap 后两者分流,重启后仍分流,DELETE 后 bob 回到 m-fast。

Co-Authored-By: ModelRouter <noreply@modelrouter.dev>
2026-10-03 07:45:54 +08:00
c19b8e6394 fix(scheduler): AUTO 遇空内容响应降级到下一个 slot(客户端曾报 "no content")
生产故障:pi 客户端报 `model "AUTO" returned a completed response with no content`,
重试延迟 8 秒。实测 AUTO 20 次有 2 次返回空 content,全部是 claude-opus-4-8。

根因(直连上游抓包确认):思考型模型先吐 reasoning_content,max_tokens 小到
思考阶段就把预算用完时,上游返回 200 / finish_reason=length,28 个 chunk 全是
reasoning_content、content 一片空白。runTier 只看 err == nil 就当成功返回,
客户端拿到一个空响应。

修复:
- resultIsEmpty:非流式路径把「无 content、无 tool_calls、无 image」的响应当作
  slot 失败继续降级。注意 ReasoningContent 不算内容——客户端要的是文本,为
  另一个模型的思考阶段扣住请求比降级更糟。
- peekStream:流式路径在出现首个真实内容前缓冲 reasoning 前导,流结束仍无内容
  则回报空结果,让 chainDrive 换 slot。缓冲只覆盖思考前导,拿到内容后立即
  转发。tool_calls delta 算内容,agent 回合不会被误判。
- 3 条判据 + 3 个变异(恒 false / 恒 true / reasoning 算内容)全部被捕获。

同时修三个 WebUI 布局缺陷(都靠截图而非 DOM 断言发现):
- 插件侧栏项只渲染图标没有标题:btn.innerHTML 只塞 pluginIconHTML(pg.icon),
  与原生页的「图标 + <span>标题</span>」不一致,侧栏是一排无名图标。
- 计费维度表 8 列挤在 465px 卡片里:table-layout:fixed 把每列压到 62px,
  23/80 个单元格溢出、数字互相重叠。改为 6 列(token 细分合并为
  「输入(新鲜+缓存)」,细分进 title)+ table-layout:auto,实测 0/60 溢出。
- 数字列 word-break:break-all 让每个字符独占一行(USD 0.56 竖排成 U/S/D),
  改 nowrap + 容器横向滚动。
2026-10-02 15:10:01 +08:00
d9652f479a fix(plugins): 插件 Lua 报错不再拖垮网关(生产事故修复)
## 事故

13:37 部署后线上 6 次 SIGSEGV 崩溃循环,8081 完全不可用,用户报大量
connect error。崩溃点固定在 internal/lua/plugins.go:invoke → L.Call →
golua StackTrace 里的 lua_getinfo。

## 根因(不是并发/GC/锁)

golua 的 callEx 在**任何** pcall 失败后无条件执行 L.StackTrace(),而
StackTrace 调 lua_getinfo,这个 LuaJIT 构建在栈够深时(带 AUTO 链轨迹的
request_end payload 正好够深)直接段错误。这是 C 层信号,Go 无法 recover,
所以一个插件的脚本错误就能带走整个进程和所有在途请求。

触发错误来自我上一轮加的 billing 日级维度:

    add(bucket(bucket(bucket(s.by_day_src, dk), payload.source)), ...)

三个 bucket( 只对应两个 ),最外层 bucket() 只收到一个参数,k=nil,于是
billing.lua:141 `tbl[k] = b` 抛 "table index is nil",**每个请求都抛**。

同时还有第二个 bug:中间层用了 bucket()(返回 emptyBucket,含 cost/requests
字段)当作嵌套容器,结构也是错的。改为 dayMap() 返回纯表。

## 修法

1. billing.lua:修正括号,多层容器改用 dayMap()。
2. **pcall 守卫**(真正的架构修复):在 setupGlobals 里注册
   __llmsproxy_call_hook,钩子改为经它调用。

       function __llmsproxy_call_hook(fn, payload)
         local ok, res = pcall(fn, payload)
         if not ok then return nil, tostring(res) end
         return res, nil
       end

   Lua 侧 pcall 在 golua 看到非零 pcall 状态之前就拦下错误,C 栈回溯路径
   永远进不去。错误变成普通返回值 (nil, msg),Go 侧记进 hook_errors 并跳过
   ——"插件出错不影响请求转发"这条承诺对脚本错误也终于成立,而不只是对 Go panic。

## 这同时修掉了那个查了很久的间歇崩溃

同一个机制解释了此前 8/20 复现、却查不出根因的 SIGSEGV(怀疑过 janitor 竞态、
GC、LuaJIT 全局状态、VM 释放时序,全部排除)。实测对比:

  TestBillingPrecedence   修复前 8/20 崩溃 → 修复后 0/20
  并发建 16 个 VM 的探针   修复前 3/3  崩溃 → 修复后 0/6
  全量 ./...              连跑 5 次全绿

那些崩溃本来就是一个 Lua 钩子错误在栈深时炸掉 StackTrace,时机随机所以看着
像并发问题。

## 判据

TestHookThatRaisesDoesNotCrashTheProcess:装一个每请求必崩的插件,连打 50 次,
断言进程存活 + 错误被记录 + 同状态里健康的 billing 插件照常工作。
3 个变异(守卫不 pcall / 守卫名写错 / 守卫未注册)全部被捕获,其中第一个直接
让 SIGSEGV 重现,说明守卫就是唯一防线。

## 线上验证

往生产插件目录放一个每请求必然报错的插件,连打 30 个真实流式请求:

  30× HTTP 200,SIGSEGV 0 次
  hook_errors 记录 count=44 且指名 zbroken-test(可观测)
  billing 照常累计(2999 请求 / $0.5668)

测试插件已移除。

回滚点:/usr/local/bin/llmsproxy.bak-real-<TS>、billing.lua.bak-real-<TS>。
2026-10-02 14:07:44 +08:00
c241a19b51 feat(stats): 用量按日/周/月/全部统计(审计文件聚合)
统计页原来只有一个视图——进程启动以来的累计。早上没人和一整周没人看起来
一模一样。加日/周/月/总四个周期。

## 口径与实现

周期视图必须走审计文件,不能走内存聚合:内存 byModel/byKey 等是终身累计,
而 recs 环形缓冲只有 500 条(defaultRingSize)。读环会把任何超过几百个
请求的周期悄悄少算——这正是要消除的那类错数。

- 日/周/月 = UTC 日历窗口(今日 / ISO 周周一 00:00 / 本月 1 日)。
  刻意不用滚动 24h:滚动窗口会让"今天"和"最近一天"边界不同,同一个数字
  随查看时刻在两张卡片间跳。UTC 也和 billing 的峰段计算同口径,峰谷小时
  不会在费用视图和用量视图里落到不同一天。
- 全部 = 复用现有 Snapshot(内存聚合),无审计文件时依然可用。
- 时间桶:日→每小时(今天内部的尖峰要看得见),周/月→每天(否则一周是
  7×24 个点、一个月 31×24)。全部视图无时间线(终身总量没有有意义的
  时间轴,硬画 500 个滚动小时点是另一种撒谎)。
- key 过滤在所有维度生效;非法 period 返回 400 而不是静默回落"全部"——
  书签里的手误应当报错,而不是悄悄换成终身数字。

## 判据(10 条 + 8 个变异全部被捕获)

窗口边界(含"周日必须回到上一个周一"这个 Go Weekday() 陷阱)、旧记录不
计入、维度独立聚合且 by_model 求和等于 total、日桶按小时且有序、key 隔离、
全部视图走终身、空窗口标记 truncated、period 校验、query 解析。

变异验证时 by_status 假绿了一次:禁用状态码聚合后判据全过,查下去是我
**根本没测 by_status**(零覆盖)。补 TestPeriodStatusDimension 后该变异
立即被捕获。判据报假问题时,先怀疑判据——这次确实是我错了。

## 真实流量核对

生产审计文件手算 vs 后端(含轮转文件):
  day   手算 3559 / 后端 3532
  week  手算 43240 / 后端 35034
  month 手算 13696 / 后端 13670
差异是核对快照与请求之间的新流量,量级一致。

一个必须说明的发现:审计文件里混着两种记录 —— Req(type/model/
prompt_tokens)和访问日志(lat_ms/status/path)。46026 行里 33450 行是
访问日志,Go 侧按 r.Type=="" 跳过。这不是 bug(CSV 导出同样如此),但
意味着任何按行数手算都必须过滤,否则会差一个数量级。

CDP 实测四周期切换:reqs 3,571 / 35,075 / 13,710 / 196,712,与后端一致,
无控制台错误,localStorage 持久化生效。
2026-10-02 12:26:26 +08:00
aa10ee8c27 fix(billing): 插件页不可达(真·空白根因)+ i18n + 溢出 + 状态页 tile
## ★ 用户报告「Billing 页还是空白」—— 上一轮的验证有漏洞
上一轮我用 goTab('billing') 直接调用验证,显示"有数据、无错误"就下了结论。
但用户是**点侧栏按钮**。真实点击路径走 goTab,而 goTab 只遍历硬编码的 TABS
常量来切换 `hidden` 类 —— 插件页不在 TABS 里,所以 #tab-billing 的 hidden
**永远不会被移除**。内容一直躺在 DOM 里(KPI/表格都填好了),只是不可见。

这个 bug 没有任何报错:注入正确、数据正确、API 200,唯一的问题是宿主页的
路由逻辑没把插件页纳入。而它是上一轮「TABS 收敛为单一常量」时留下的:
收敛让三处共用一个常量,却没让插件页进入它。

修法:goTab 同时遍历 PLUGIN_PAGES。PLUGIN_PAGES 从 const 改为 var 并**提前到
goTab 之前声明** —— const 在文件后部声明的话,goTab 的读取落在 TDZ 里,
第一次点击插件页就会抛 ReferenceError(同类问题这个文件里已是第二次)。

判据 TestPluginPagesAreReachableByGoTab 锁两件事:goTab 遍历插件页集合 +
声明在使用之前。变异验证:删掉遍历 → 红;var 改回 const → 红。

## 插件 UI 不跟随多语言
宿主的 applyI18n/data-i 只覆盖**宿主渲染的标记**;插件注入的 HTML 对它不可见,
所以整个 UI 切中文时 Billing 页还是英文。

pluginAPI 增加 lang(getter,实时值)与 onLangChange(切换回调)。
billing 页所有文案改走双语字典:KPI、表头(fresh/cache/cache%)、区块标题
(占位后由脚本填)、空态、状态页 tile 标签。切换时立即重渲染标题,
不用等下一次 fetch。

## 部分页面超出 UI 区域
#main 只有 overflow-y,插件页内容(8 列表格 min-width、长字符串)会横向撑破。
两层修:插件 pane 统一 min-width:0/max-width:100%/overflow-x:auto(第三方
任意 HTML 的兜底,与原生 pane 一致);billing 的宽表格在自身容器内滚动。

## 状态页 tile 的 TypeError(每次重绘都报)
tile 的 tick() 在 await 之后直接 getElementById(...).textContent = ...,
但状态页每次刷新都整体重建 pane,元素可能已不存在 → null 属性赋值。
await 之后重新取元素并判空。

## 顺手补的缺口
上一轮加了缓存表格列,但 KPI 卡片漏了(那次替换 assert 失败后重试只重做了
表格)—— 缓存命中率在表格里有、KPI 里没有。本次补上。

## 验证
真实浏览器(禁缓存、真实点击侧栏按钮):pane 可见、KPI 9 项、表头双语、
语言双向切换正确(Per source ⇄ 按源)、无水平溢出、无 billing 控制台错误。
生产数据:Total USD 0.566798 / 732 请求 / 降级 231 / 2.09 亿 prompt tokens。
387+ 测试全绿。

## DSL(进行中,未完)
config.BillingDSL(active + profiles + rules,rule 按 url 匹配 mode=free/
token/subscription/unpriced)与 internal/billing.Compile(url 规则 → 插件
prices 表,含峰谷窗口的形状编译 —— 之前手写 JSON 两次弄错的正是这个形状)
已落地并通过校验/编译;core 启动接线已写。profile 切换 API 与 WebUI 选择器
未做,生产 config.yaml 也尚未写 billing 段 —— 下一轮继续。
2026-10-02 11:47:24 +08:00
fbdf0dea10 fix(billing): 缓存命中统计缺失 + Billing 页空白 + 侧栏图标
三个问题都来自生产实测,不是代码审阅。

## 1. 缓存命中被计费却不被统计
网关确实从上游 usage 提取了 prompt_cache_hit_tokens(审计里能看到
cache_hit_tokens: 270104 / cache_reported: true,占 prompt 的 99.9%),
costFor() 也用它给缓存段定价了 —— 但**没有任何 bucket 记录它**。
结果:一个 99.88% 命中率的网关,报表显示 prompt_tokens 却看不出其中
多少是缓存读,也无从按源/模型/key 看命中率。

每个 bucket 现在多三个字段:
  cache_hit_tokens    命中数(按上游上报)
  cache_fresh_tokens  未命中的 prompt
  cache_reported_reqs 上游确实上报了缓存数的请求数

第三个字段是刻意的:**「零命中」与「上游根本不上报」在命中总量里完全一样**,
而它们在「缓存折扣有没有生效」这个问题上含义相反。没有它就无法区分,
只能猜。

chat.go 的 payload 之前**没有** cache_reported(审计有、插件没有),
所以任何插件侧的缓存统计都只能猜 —— 已补上。

旧 state 文件的 bucket 没有这些字段:Lua 里 nil + number 会抛错,而钩子抛错
会让**该请求完全不记账**(一个统计缺口会变成静默缺口)。add() 里做了回填。

UI 增加 fresh/cache/cache% 三列 + Cache hit rate KPI;未上报的显示 n/r 而不是 0%。

## 2. Billing 页空白:render() 引用了未定义的 s
`render(st)` 里两处 KPI 写成 `s.degraded_reqs`,ReferenceError 让整个渲染
中断,所有表格停在初始的空 innerHTML。症状是「页面加载了但什么都没有」,
而 /api/plugins/billing/state 返回 200 且有真实数据 —— 载荷完全正确,
DOM 是空的。

更糟的是 refresh() 里的 `catch (e) { /* never break the page */ }` 把错误
**静默吞掉**了:网络面板一切正常,页面什么都没有。现在 catch 会
console.error(仍然不抛,装饰性组件不该拖垮宿主页,但必须留痕)。

## 3. 侧栏图标
billing 声明 icon = "💰",而原生 tab 全是内联 SVG(stroke: currentColor)。
emoji 尺寸不对、不跟随主题。

WebUI 增加 pluginIconHTML:插件图标可以是文本,也可以是内联 SVG。
**SVG 走严格白名单**(tag + 属性都是 allowlist,不是 denylist)——
插件是在运维者浏览器里跑的第三方代码,不能"信任插件";但也不能直接拒绝
SVG,因为那是唯一能和原生 tab 视觉一致的方式。

用真实 Chromium 验证 12 个用例,全部挡住,包括 foreignObject 里嵌 HTML
命名空间 <img onerror> 这个经典绕过(整体丢弃,所以 img/onerror 也没了)。
★ node 里没有 DOMParser/jsdom,所以没法在单测里跑这个过滤器 —— 用正则近似
会得到一个"测试通过但浏览器里失效"的过滤器,这比没有测试更糟。

顺带修了过滤器的两个真缺陷:输出里嵌套了空 `<svg></svg>`,且 viewBox
是从包装元素读的(永远是 null)而不是插件自己的,所以任何自定义 viewBox
的图标都会丢失。

## 判据(新增 7 项,全部变异验证)
写「注入脚本能否正常执行」这个守卫时我错了四次:
  1. 静态扫「已声明的名字」→ 把 HTML 字符串里的 CSS 类名(class/div/td)
     全报成未定义
  2. 用 CSS 选择器解析器查样式表 → 报样式表本身坏了
  3. 只挂 process 的 uncaughtException → 脚本在 IIFE 里异步跑,错误是
     unhandledRejection,判据对原 bug 全绿
  4. 只查「有没有抛错」→ render() 开头是 `if (!st) return`,传错字段是
     **静默 no-op**:不抛、不打日志、不报错,只是页面空白
最终判据是:在 node 里用 DOM stub 真跑一遍,同时要求「无异常」且
「至少写进一个容器」,并监听 console.error。变异验证:还原 s → 红;
render 收到 undefined 字段 → 红。

表头/行列数一致性也有守卫:row() 加了缓存列而表头没加时,表格会整体错位
(cache% 落到 completion 列下)—— 渲染正常、有数据、但要仔细看才发现。

## 生产验证
重启后价目表与累计账完整保留(1.17 亿 prompt tokens)。
新请求缓存统计生效:cache_hit 947,436 / cache_fresh 888,
cache_reported_reqs 7 / 395(其余来自旧 state,正是该字段存在的意义)。
真实浏览器:表格 3 行、KPI 7 项、表头 name/cost/reqs/prompt/fresh/cache/cache%/completion、
SVG 图标 currentColor 渲染、控制台无 billing 错误。391 个测试全绿。

## 另发现一个无关 bug(未修)
首页 stats 图表抛 IndexSizeError: arc 半径为负(-2),在 ui/index.html 的
paintStats 附近。属状态页图表,不在本次范围。
2026-10-02 11:16:15 +08:00
8de1499c40 fix(ui): 插件元素注入被宿主页面重建擦除 —— 元素型注入从未真正生效
## 现象
部署示例后打开 WebUI:侧栏有 Billing 页,但**状态页上没有任何计费组件**。
插件明明声明了 elements,/api/ui-inject 也确实返回了 mount(1570 字节)。

## 根因(不是缺功能)
七个宿主页面的渲染函数都用 `pane.innerHTML = ...` **整体替换**自己的 DOM。
`renderStatus` 在 `injectPluginUI()` 之后由 refresh() 立刻调用,于是刚挂上的
plugin-el 连同整个 pane 一起被下一次赋值销毁。

时序上它**从没有过"显示一帧"的机会**:注入 → refresh("status") → innerHTML 覆盖。
所以症状是"元素从来没出现过",而不是"刷新后消失"——这正是我先前据
/api/ui-inject 返回值判定"注入正常"而漏掉的地方:**载荷到达 ≠ DOM 存活**。

Billing 页不受影响,因为它属于 PLUGIN_PAGES,走插件自有 DOM,不经宿主重建。
于是看起来像"页面注入有效、元素注入无效",把排查引向插件声明本身。

## 修法
- mountPluginElements 改成具名可重入函数,并注册进 PLUGIN_MOUNT_HOOKS
- refresh() 在**唯一出口**统一调 remountPluginElements(),而不是给七个渲染函数
  各加一次调用——后者是多一处会忘的地方,而忘记的后果是静默的
- 每个挂载点按 data-idx 幂等:宿主重绘时若该 pane 已有该元素就直接返回,
  否则插件的 <script> 会每次重绘都跑一遍,计数器静默翻倍

## ★ 验证方式换了:真实浏览器,而不是 payload
静态测试和 curl 都看不出这个 bug(载荷完全正确)。用 CDP 连本机共享浏览器实测:

  修复后:首屏 tile=1,页面重建后=1,连续重建 5 次仍=1,console 无错误
  回退后:tile 全程=0

  对照二进制(把 done 改成空函数重编译)实测首屏就是 0,
  **证明"从未显示过",不是"显示后消失"**。

## 判据与变异
TestPluginElementsSurviveHostRebuild 锁住:remountPluginElements 存在、
PLUGIN_MOUNT_HOOKS 在使用之前声明(const TDZ 会让首屏直接抛错)、refresh 挂了重挂、
挂载按 data-idx 幂等。

三个变异全部被抓住:撤掉 refresh 的重挂 / 去掉幂等守卫 / 把 const 声明移到 push 之后。
★ 第一次跑第三个变异时**判据正确地没报**,因为我的替换脚本命中了注释里的同名文本,
真正的 const 没被移动——是变异无效,不是判据有洞。换按行定位后如期变红。

## 同时补上部署示例
packaging/config.example.yaml 里补 plugin_dir 说明(之前只有 online 部署路径踩过)。
实测升级路径本身是好的:给已有配置加 plugin_dir 后,首次启动会自动 seed 内置
billing 插件,无需手工放置文件。
2026-10-02 09:04:47 +08:00
1c690611f8 feat(gui): WebUI 与 Electron 壳的插件安装/删除/禁用/编辑
## WebUI:新增「插件」页
- 列表来自 on_disk(不是 loaded 集合)——**加载失败的插件也必须显示并带错误**,
  否则一个语法错误看起来和"插件没装"完全一样
- 启用/禁用(PUT {"enabled":bool})、删除、编辑源码、安装/覆盖
- 显示 hook_errors:插件抛异常在别处毫无痕迹,没有这一栏的症状就是
  "功能就是不work"
- 插到 dropzone 与代码编辑器都做了泛型化(bindDropzone / openCodeModal),
  适配器与插件共用一份,而不是复制第二份只改 4 个 id 的函数

## TABS 收敛为单一常量
tab 清单原本是字面量散在三处:goTab、refresh()、admin-only 隐藏列表。
加一个 tab 意味着三处都要记得改,漏一处就是"路由认得但界面不显示"——
和今天早些时候 chain_step 漏报同一类静默缺口。现在只有 const TABS。

## Electron 壳:设置面板里的插件管理
渲染进程不能直连内嵌核心(没有 key、不知道端口),所以走 IPC:
  renderer → plugins:proxy → main → HTTP /api/plugins
代理是 (method, path, body) 透传而不是固定命令表:固定表每加一个端点就要扩,
而"按钮存在但什么都不做"比"没有这个按钮"更糟。透传让渲染层能调用核心将来
新增的任何 /api/plugins 路由,路径在主进程校验。

## ★ GUI 此前零测试,而本次改动就引入了三类"看起来没事"的问题
1. 引用了不存在的 CSS 类(.tag / .sm)——渲染成无样式文本
2. 引用了不存在的 helper(esc / escAttr)——那是 WebUI 的,renderer/app.js
   是独立文档,点击时 ReferenceError
3. .ghost/.primary 只在 .form .actions 作用域内生效,插件按钮在 .pl-acts 里
   于是是无样式裸按钮

补 4 个静态判据(不启动 Electron,守卫的正是"打开应用才看得见"那一类):
  TestGUICSSClassesExist          用到的类必须在样式表里定义
  TestGUIHelperFunctionsAreDefined 被调用的函数必须有定义
  TestGUIPluginPanelIsReachable  面板在 overlay 内、按钮已绑定、打开设置会加载
  TestGUIIPCPathIsConstrained    代理必须限定 /api/plugins 前缀并拒绝路径穿越

写第一个判据时我错了三次:CSS 解析器先丢最后一个 selector、再把变量块当
selector、最后漏掉复合选择器(.tb-btn.tb-close)。两次"判据自己坏了"的
教训和本项目一贯一致——**判据出错的信号是它报了一个假问题**。现在改用宽松的
token 提取 + 显式的 guiKnownUnstyled 豁免表(blob/tgl/rail 是既有无样式类,
不是本次引入,失败它们只会让判据对新工作失去意义)。

## 变异验证
  改坏唯一的 CSS 定义(.pl-empty)→ TestGUICSSClassesExist 红
  改坏 helper 名 → TestGUIHelperFunctionsAreDefined 红
★ 第一次变异我改了 .pl-broken,判据**正确地没报**——因为它还被另一条规则定义。
  这是变异选错目标,不是判据有洞;换 .pl-empty 后如期变红。

363 个测试全绿。
2026-10-02 08:47:08 +08:00
a51a6811a6 feat(plugin): Lua 插件机制 + 计费插件 + 插件文档
插件 = plugin_dir 下的单个 .lua 文件,做两件事:挂请求流水线的钩子、在启动时
贡献 WebUI 界面(整页或往现有页面追加组件)。两者独立。

## 流水线 stage(三个)
  request_start  已解析鉴权、未选源
  routed         已选定 (source, model)、未发往上游
  request_end    每请求恰好一次,带最终计量
request_end 挂在 gateway.writeRec——四条入口路径(直连/AUTO × 流式/非流式)的
唯一汇合点:既不漏(流式 token 只有流结束才知道)也不重。

## 计费插件(plugins/billing.lua,默认 seed,开箱可用)
源 / 模型 / 密钥三个维度定价。token 价优先级 keys > models > default;per_request
固定价是**叠加**的(生图模型可以既算 token 又收固定费)。单位是 USD/单 token,
即各家 provider 的公布口径。累计 total / by_source / by_model / by_key / by_day。
失败请求保留 token 费用、丢弃固定费(可经 count_failures 翻转)。
界面 = 一个独立页 + 状态页顶部一块总开销 tile。

## 一个明确的设计边界
计费插件**只报表,不执法**。网关自己的配额会计(stats.go,入口强制)才是限额
权威,插件不参与任何路由/配额决策。两套独立会计若对不上,比一套功能略少的
更糟。

## ★ 中途改掉的一个根本设计错误
最初让插件复用适配器的**弹性 worker 池**(多状态)。这对适配器是对的(它们无
状态),对插件是错的:计费插件往 plugin.state 累加,多状态意味着总量被劈成
几份;而 SetState 写价格只写进其中一个 worker,钩子恰好跑到另一个时**所有请求
按 0 计费**。改为**单状态 + 互斥锁**。代价写进文档:钩子必须短、同步、不阻塞,
卡住的钩子会卡住所有插件的钩子。
这个 bug 是测试逼出来的——先写了 SetState+Fire 的用例,数字全是 0 才挖出来。

另一个连带缺陷:只带 prices 的 PUT 会整体替换 state,把累计量清零。改为
prices/state 分离——prices 是配置、state 是历史,改价不动账。

## 撞到的三个 Lua 绑定的坑(都写进注释)
  - SetGlobal **会 pop 栈**:连着调两次,第二次从空栈取,赋成 nil
  - GetField 索引越界是 **SIGABRT 整个进程**,不是 panic,recover 救不了
  - Call(nargs, n) **不接受函数索引**,它调的是 nargs 个参数正下方那个;
    传索引会调到参数上("attempt to call a table value")
另外 GetField/SetField 用绝对索引,SetTop(0) 之后必须重取。

## 错误隔离
钩子 error() 不影响转发:捕获 → 记进 hook_errors → 跳下一个插件。适配器出错
会让源进冷却,插件出错**零惩罚**——插件是可选功能。/api/plugins 的 hook_errors
让"坏掉的插件"可见而不是静默消失。

## 界面注入
GET /api/ui-inject 一次返回所有插件的扩展(侧栏需要全部 page 才能建好)。
WebUI 在首次 render **之前** await 注入:先插 HTML 再重建 <script> 让它执行
(innerHTML/template 插入的 script 不会执行,这正是要的效果——避免脚本跑在
自己 DOM 之前)。注入失败不影响仪表盘。
browser 侧 pluginAPI 暴露 fetchState / postState / onTabShown。

## 文档
docs/plugins.md —— 快速上手、加载与热更新、三个 stage 的完整字段表、界面扩展、
状态与 HTTP API、运行时约束(单状态/异常隔离/内置函数)、计费插件的定价与
计费策略、排错表、与适配器的对比表。

## 判据(328 个测试全绿,插件相关 33 个)
  - 计费断言的是**具体金额**(0.00625 / 0.0402 / 0.0075…),不是"能加载"
  - 4 个变异都红:钩子异常不隔离 / prices 清空累计 / 忽略 key 优先级 /
    毫秒时间戳不换算
  - UI 侧 6 个判据把注入顺序、script 执行时机、pluginAPI 名称、tab 路由、
    anchor 四种形式、失败非致命全钉住
  - 鉴权:state 读任意角色、写仅 admin
2026-10-02 00:37:29 +08:00
18e422a943 fix(sources): 源编辑不再清零 proxy_url / api_key_env / timeout
bf0657b 修好了 api_key,但 upsert 仍会重写整个源,于是请求无法表达的字段
一律被重置为零值。这四个字段的后果都不是"少个配置项":

  - api_key_env 丢失 ⇒ 盘上无明文密钥的源变成无凭据源,写操作返回 200,
    下一次调用上游才 401。而 README 恰恰把这个特性当作卖点在宣传。
  - proxy_url 丢失 ⇒ 一个走代理的上游变成直连(或反之),且完全无声。
  - timeout / queue_timeout 丢失 ⇒ 退回默认 120s / 60s。

触发路径不是只有脚本:WebUI 的 saveSource() 发的 payload 只含表单上的
11 个字段,而 editSource() 表单里根本没有这 4 项 ⇒ 运维在界面上改个并发数
就会静默清掉它们。

修法用「存在性」语义而不是「空即继承」:

  - 不传   → 保留已存值(部分更新的客户端要的就是这个)
  - 传了   → 覆盖,包括传空串表示清空

api_key 刻意保留它原有的「空即继承」规则,不跟着改成指针:该规则已随
v1.7.6 发布,脚本依赖它;而凭据丢失比代理丢失严重得多。两个字段的失败
模式相反,所以规则相反——这一条写进了两处注释。

另一处是差一点的:config.Source 把 Timeout/QueueTimeout 标成 json:"-",
所以 reveal 接口的结构体序列化**根本不返回它们**。表单读不到 → 输入框
恒空 → 而输入框每次都回传 → 每存一次就把 timeout 清零。等于把刚修好的
丢字段换个方向又造了一个。因此 reveal 分支现在显式返回 duration 字符串。

判据:
  - TestWebUIEditPayloadPreservesRoutingFields 用的是 WebUI 真实 payload
    的逐字节副本,并同时断言"确实改动的字段生效",否则"什么都不写"也能过
  - TestSourceEditPreservesAPIKeyEnv 单独盯 api_key_env(唯一造成凭据丢失的)
  - CanBeSet / CanBeCleared 分别盯两个方向:只有"空即继承"的实现过不了
    CanBeCleared(清空代理框会永远保留旧代理)
  - 持久化判据**重新加载 config.yaml 并按语义比对**:300s 会被重新序列化成
    5m0s,按字符串匹配是假红(我自己先踩了一次)
  - TestSourcePayloadCoversEveryEditableField 用反射卡住"这一类":新增
    Source 字段而没接到 API 上时立刻变红。反射查结构体而非 marshal 结果,
    因为指针 + omitempty 会合法地从序列化输出里消失,那正是"未提及"信号
  - 三个 UI 契约判据把 JS 侧也钉住(表单必须回传、必须从 reveal 读)

变异验证(每次都先确认 build 通过,再数红格):
  1. 去掉覆盖逻辑        → 7 个判据红
  2. 改成"空即继承"      → CanBeCleared + ClearIsScoped 红
  3. reveal 不返回 duration → TestSourceRevealExposesDurations 红
  4. 表单不回传 api_key_env → TestUIEditFormRoundTrips... 红

顺带修正 /api/v1 索引:DELETE /api/keys 的路径段写的是 {name},实际是 key
本身;PUT /api/keys/{key} 实现了却没列。

全量 + vet + race 全绿;WebUI 内联脚本过 node --check。
2026-10-01 20:36:48 +08:00
0121d23f91 fix(tokens): 流式统计改用上游真实 usage,图片不再记 token
两处 token 单位错误,均影响 per-model 配额计费:

1. 流式路径的 prompt/completion 只是「字节÷3」估算。
   pumpStream 明明收到了上游最后一帧的真实 usage,却只发给客户端、
   从不回写审计记录,于是配额按估算值扣。生产实测同一请求:
   上游 prompt=37/completion=179 → 记账 27/262,prompt 低估 1.4x、
   completion 高估 1.5x(双向失真)。同模型流式 completion 中位数
   是非流式的 4-27 倍。非流式路径本就用真实值,两路不一致。
   修法:lastUsage 非零时写回 rec.Prompt/rec.Compl,估算降为兜底
   (上游不报 usage 时仍保留原估算行为)。

2. 图片请求把「图片张数」记成 completion_tokens。
   rec.Compl = int64(len(resp.ImageData)),len 是切片长度即张数
   (生产 38 条 image 记录全是 1),且被计入 token 总量。
   图片生成无 token 概念 ⇒ 新增 Req.ImageCount 独立字段,
   Prompt/Compl 归 0;UI 记录表 image 行改显示张数(新增 i18n thImgs)。

顺带补 TestUILocaleKeyParity:此前无人校验 zh/en 键集合一致,
单边加键不会报错,只会显示原始键名。

新增 token_units_test.go(定值上游 6 项),做过变异验证:
回退修复实测复现 stream=16/173 vs chat=44/100、image completion=3。
2026-09-28 22:19:12 +08:00
c51066f0b6 refactor(quota): 配额改为按模型,删除整钥总配额
用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
2026-09-27 19:02:13 +08:00
f814ff7468 fix(webui): 修 7 处弹窗关闭错对象 + 模板管理器变量遮蔽
上一提交只修了自己新加的两处弹窗,全站其余 7 处是同一缺陷:所有对话框
共用 id="modal-wrap"(CSS `#modal-wrap:not(:empty){display:flex}`)且可以叠
加(seed-key 提示就盖在密钥页上),而
`const w = $("#modal-wrap"); w.remove()` 移除的是**文档里第一个**,不是用户
刚提交的那一个。

逐处改为两种安全写法:
- 能拿到按钮的(saveSource / saveTemplate / sortScopeSave / scopeSave /
  keyQuotaSave / createKey / downloadStatsCsv / downloadKeysCsv /
  scrAddFromForm):`btn.closest("#modal-wrap")`
- 拿不到按钮的:新增 `closeTopModal()` 取**最后一个**(用户看到的那个),
  并作为所有 `if (w) w.remove()` 之后的兜底
- 顺带把 sortScopeSave / downloadStatsCsv / downloadKeysCsv / scrAddFromForm
  的签名补上 btn / this 参数 —— 否则 .closest 恒为 null,表单永远不关

同时修一个相邻的既有 bug:`openTemplateModal` 的 `.map((t) => ...)` 用 t 做
循环变量,模板字面量里又调 t("srcEdit"),t 被遮蔽成对象 ⇒ 打开模板管理器
直接抛 `t is not a function`,整个弹窗渲染失败(main 上就有,git show 确认)。
参数改名 tpl。修后模板管理器完整渲染(浏览器实测:DeepSeek / 智谱 / Kimi /
SiliconFlow 各行 + Edit/Delete 按钮文案全部正常,零异常)。

判据从 2 处扩到全量(internal/gateway/ui_quota_contract_test.go,+3 例):
- 全文档扫描:任何 `.remove()` 配裸 `$("#modal-wrap")` 即失败
- 9 个关闭对话框的处理器必须走 .closest 或 closeTopModal
- closeTopModal 必须取 all.length - 1(首尾颠倒就是原 bug)
- 用 .closest 的处理器,其签名必须真的有 btn 参数 —— 否则查找恒为 null,
  表单永远不关
5/5 变异全被抓:sortScopeSave 去 btn 参数、closeTopModal 取第一个、scopeSave
退回裸选择器、saveSource 退回裸选择器、downloadKeysCsv 删兜底。

浏览器实测(共享 Chromium CDP,真实进程,每处都插入一个「decoy」弹窗
占据文档首位,复现原 bug 的触发条件):
- scopeSave / keyQuotaSave / scrAddFromForm / saveSource / saveTemplate
  五个处理器:自己的表单关、decoy 保留 ✓
- 零 JS 异常
★ 测试自身踩了两个坑,都不是代码问题:① `document.querySelector(sel) && .click()`
  在 CDP 里求值为 undefined,改成箭头函数;② 编辑模板时没填名字就点保存,
  saveTemplate 因 `if (!nm)` 早退、fetch 零调用 —— 一开始我把这个误读成
  「修复失效」,加 fetch 拦截 + 读 #s-name 的值才定位到是测试数据缺失。
  ⇒ 「点按钮没反应」要先分清是「事件没触发」「请求失败」还是「早退」。

(cherry picked from commit b5c3fea0bb)
2026-09-27 18:44:41 +08:00
ef631b43dd feat(webui): 密钥配额表单 + 修复弹窗关闭错对象
功能:让 per-key 配额在 WebUI 里可配置可见,之前的实现只有 API 与
config.yaml 能配。

- 密钥卡片头部显示配额徽标(token / 请求数 + 重置窗口),admin key
  不显示编辑入口(服务端本就永不受限,给入口只会让人以为配了会生效)。
- 新增「配额」编辑弹窗:token 配额、请求数配额、重置周期(复用既有
  的 period 词表与 n-hour 联动),预填从 canvas 的 data-* 读。
- 创建密钥弹窗同步加配额字段;选 admin 角色时自动禁用(同样因为服务端
  忽略 admin 的配额)。
- 「我的密钥」页新增 KEY-WIDE QUOTA 列,用户能看到自己这把 key 的预算。

修一个真 bug:保存弹窗用 $("#modal-wrap") 关闭自己,而全站弹窗共用这个
id、且可以叠加(seed key 提示就盖在密钥页上)。实测(共享 Chromium
CDP,seed 提示与配额弹窗共存)确认:保存后被移除的是 seed 提示,配额表单
反而留在屏幕上 —— 症状是「保存了但弹窗没关」,指向的方向完全错。改为用
点击的按钮 btn.closest("#modal-wrap") 解析自己的弹窗。createKey 有同样
问题,一并修。既有文件里另有 7 处同样写法,未动(不在本次范围,且新判据
只对本次改的两处断言,避免误伤)。

判据新增 internal/gateway/ui_quota_contract_test.go(6 例):
- 两个表单必须用 .closest 解析自己的弹窗
- putScope 必须带上 4 个配额字段(API 视其为指针,省略=清空预算)
- 创建请求必须真的发出配额字段
- **数据流判据**:徽标要真读 k.token_quota 等、编辑表单要真读
  canvas 写的 data-kquota 等。只查字面量存在会漏 —— 字段躺在死分支里
  判据照样通过(这是本轮实际踩到的:keyCapBadges 经 keyPeriodSuffix
  间接读 k.period,被判据抓到后我把读取显式化而不是放宽判据)
- 弹窗扫描先剥注释,否则修复说明里引用的字面量会被当成违规
- 复用既有 ui_contract_test.go 的 jsFunctionBody(大括号配平);
  自己第一版用 2000 字符固定窗口,被长注释顶开后仍在窗口外命中后面
  函数的同名字段,读起来像通过 —— 窗口法在这里是假判据

7 个变异全部被抓(unsafe 关闭、putScope 丢字段、createKey 丢字段、
徽标不读字段、canvas 不写 data-*、kq-hours 改名、周期词表缺项)。

浏览器实测(共享 Chromium CDP,真实进程 + 加密配置):
- 徽标渲染 1.0K·1h / 5×·1h;编辑框预填 1000/5/hour,hours 框按周期联动
- 保存后回读 250000/77/nhour/6,徽标更新为 250.0K·6h,toast Saved
- 零 JS 异常
- **关键回归**:编辑模型 scope 后配额仍是 250000/77/nhour,未被清空
- 创建带配额的 key,服务端确认 {t:50000,r:300,p:week,role:user}
- user 视角「我的密钥」显示 777·1h 与 9×·1h

文档:README.md / README_EN.md 补「密钥用量配额」小节(配置示例、
周期词表、429 语义、admin 豁免、整点分桶最晚晚 1 小时释放、PUT 的
省略 vs 0 语义、429 响应样例),特性列表各加一条。

(cherry picked from commit ce66c7f6c2)
2026-09-27 18:44:41 +08:00
f6baa13583 feat: WebUI 改为依赖 /api/v1,UI 与 agent 共用一套 API 契约
- sources / sort / keys 三个页面的数据源从 /api/sources 切到 /api/v1/sources
  (写操作仍走 /api/sources:v1 是只读门面,不做变更)
- 编辑弹窗改用 /api/v1/sources/{name}?reveal=credentials(admin-only)取明文 key。
  这是必须的:表单要整体回传源,若不回填 key,改个端口就会把 key 清空。
- 遮蔽视图仍是默认,只有显式 reveal 才返回明文

端到端验证(真浏览器 + 临时实例,非仅 API 测试):
- sources/sort/keys 三页实际发出 GET /api/v1/sources,0 console error
- editSource('demo') → reveal=credentials,#s-key 与 #s-url 正确回填
- 写入往返:改 base_url /v1→/v2 后重开,key 仍在(未被清空)
- 落盘 api_key 明文残留 0、密文 1

测试:+1(reveal 必须 admin,否则任意 user key 可读全部凭据)
变异验证:reveal 去掉 admin 校验 → 403 断言变红

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 926b9f6565)
2026-09-27 17:12:03 +08:00
c309448414 feat(webui): Mono theme — pure white in light mode, pure black in dark mode
Adds a fourth accent alongside sakura/ocean/violet. Unlike those it is not just a
different hue: the coloured themes are glass surfaces (translucent cards with
backdrop-filter) floating over an animated gradient-mesh background, so setting
--card:#ffffff there still renders as a tinted grey. Mono therefore also switches
off the translucency and hides the blobs, so #ffffff is actually #ffffff and
#000000 is actually #000000, with greys carrying the hierarchy that hue carries
elsewhere. A side effect worth having: no backdrop-filter and no animated blobs
makes it the cheapest theme to render, which helps on weak GPUs and over remote
desktops.

Both light and dark variable blocks are defined, so the existing light/dark
toggle drives it with no extra wiring: light -> white, dark -> black.

Also fixes a latent theme bug found while checking contrast on black: the active
chart's grid baseline assigned the literal string "var(--line)" to
ctx.strokeStyle. Canvas 2D does not resolve CSS custom properties, so that was an
invalid colour the browser ignored, leaving the previous fillStyle (black) — an
invisible baseline on every dark theme. Colours used on a canvas now go through a
cssVar() helper.

Tests: TestUIThemeMatrix asserts every accent defines BOTH a light and a dark
block plus a picker button (a half-defined theme shows up as unreadable text, not
as an error); TestUIMonoThemeIsFlat pins the opaque surfaces and disabled blobs;
TestUICanvasColorsResolveVars fails if any ctx.strokeStyle/fillStyle is handed a
raw var().

Unrelated packaging fix in the same commit: dist:linux only built deb+AppImage
while build.linux.target listed rpm too, so `make gui-dist` silently skipped the
rpm that release builds are expected to produce. Makefile/README wording updated
to match.
2026-08-30 10:04:53 +08:00
3ddae41f0c fix(gateway): aggregate the full audit history, repair record paging
Two problems reported after the on-demand log work landed.

1. Dashboard totals were wrong. LoadAudit only replayed the last 4 MB of the
   audit file, so requests/tokens/per-key rows reflected a window instead of all
   time — a regression in reported numbers, not just in presentation.

   The aggregates are now built by streaming EVERY audit file (oldest first, so
   the hourly quota buckets keep their intended trailing window) and keeping
   nothing per record: aggregate maps are keyed by key/model/source, so their
   size is bounded by cardinality. Measured on the production host: 29 MB /
   221k lines / 37k requests in ~260 ms at startup.

   What stays bounded is the RAW-record ring: a fixed-size reqRing keeps only the
   newest maxRecs records, so the ~25 MB that used to be spent appending every
   record into a slice is still saved. auditReplayBytes is gone, and
   replayPartial now means "an audit file could not be read", which is the only
   remaining way for the totals to be incomplete.

2. Scrolling to the bottom stopped loading more records. Two independent causes:

   * paintRecords rebuilt the entire table on every 5s poll whenever the row
     count did not exceed the first screen — the "is a paged view live?" test
     compared row counts and matched exactly on the first refresh — wiping loaded
     pages and resetting scroll position.
   * IntersectionObserver only fires on TRANSITIONS. With a short list, or after a
     page whose rows all duplicated the first screen, the sentinel stayed visible
     and never fired again.

   paintRecords now builds once (recsState.built) and later polls PREPEND only
   genuinely new rows; attachRecsObserver adds a scroll-position fallback;
   fillRecordsViewport loads until the list actually overflows; and
   loadMoreRecords chains (bounded) when a page yields no new rows, since the
   first fetch necessarily overlaps the first screen.

TestUIRecordsPagingWiring pins all four mechanisms structurally, since none of
them is reachable from Go. Test names/comments referring to bounded replay are
updated to describe the bounded RING instead, and both READMEs now state that
totals come from the full history while records are paged.
2026-08-30 09:29:42 +08:00
882288f67f fix(webui): send DELETE when removing keys and adapters
Deleting a gateway key from the admin UI did nothing and reported
"use GET /api/keys": delKey() called api() with an empty options object, so
fetch defaulted to GET and the request landed in the GET branch of
handleKeysAPI. delAdapter() had the identical bug and reported
"adapter code not exposed; edit in UI".

This is the third instance of the same mistake — ab20f1b fixed delSource and
delTemplate, missing these two — so it is now pinned by tests instead of by
review:

  * TestUIAPICallsDeclareMethod walks every api() call in the embedded
    index.html and fails if one passes an options object without a method
    (an AbortSignal-only read is allowed, being a deliberate GET).
  * TestUIDeleteHelpersUseDelete / TestUIMutatingHelpersUseWriteMethods pin the
    verb of each removal and write helper by name.
  * TestKeyDeleteRoundTrip covers create -> DELETE -> gone -> second DELETE is a
    clean 404, and TestCannotDeleteOwnKey keeps the lockout guard.

The 404 bodies for GET /api/keys/{key} and GET /api/adapters/{name} now name the
verb to use ("DELETE /api/keys/{key} to remove"), because that message is what a
mis-methoded client actually shows its user; "use GET /api/keys" read as though
the caller had done nothing wrong.

delSource's indentation, broken by ab20f1b, is also straightened out.
2026-08-30 09:07:05 +08:00
6575058556 feat(webui): scroll-paged record table, pool metrics, probe-aware health tags
Front-end half of the on-demand log loading plus the observability for the two
new scheduling mechanisms.

Records table:
  * the first screen comes from the dashboard poll, and an IntersectionObserver
    sentinel below the last row pulls the next page from /api/stats/records as
    the user scrolls.
  * rendered rows are capped at RECS_MAX_DOM=1000 (oldest rendered rows are
    dropped) so a long scroll cannot grow the DOM without bound, with a "back
    to newest" button to reset cheaply.
  * releaseRecords() drops the buffer, disconnects the observer and aborts the
    in-flight fetch (AbortController) on tab switch, on key-filter change and
    on pagehide/beforeunload — leaving the page releases everything at once.
  * the poll no longer rebuilds the table once extra pages are loaded, so the
    5s refresh cannot throw away scrolled history.
  * when the server reports replay_partial, the filter row states that the
    aggregates cover the recent audit tail and points at CSV for full history.

Adapters page shows each pool as created/max plus in_use/idle and the live
grow/shrink steps, with the sizing rule in the hover text.

Priority page health tags now distinguish cooling / probe-ready / probing
instead of a flat "cooling", and the tooltip spells out when the window opened,
when the single probe is allowed through and when the slot clears completely.
core.AutoSlotState carries cooldown_from / probe_after / probing / probeable to
feed this.
2026-08-30 08:06:11 +08:00
ab20f1be48 fix(webui): specify DELETE method for source/template removal
delSource and delTemplate called api() without a method, so fetch
defaulted to GET; the DELETE handlers never ran and the UI silently
left the item in place.
2026-08-27 12:14:41 +08:00
8334cffbc9 feat(templates): source templates for multi-key balancing
A template stores every source field except name and api_key, so
operators spin up N key-bearing sources from one shared skeleton
instead of duplicating the whole source block N times.

- config: SourceTemplate type + RuntimeConfig.SourceTemplates stored
  in runtime.json alongside runtime sources
- store: UpsertTemplate / ListTemplates / RemoveTemplate
- core: Templates / SaveTemplate / RemoveTemplate
- gateway: GET/POST/DELETE /api/source_templates
- webui: source list gains a Templates button opening a manager with
  per-template edit/delete; the add-source dialog gains 'from
  template' (event-delegated picker) and 'as template' (card modal)
  buttons in its header; z-index fixed so the template editor layers
  above the manager
2026-08-27 12:09:38 +08:00
a42ff62d06 feat(scheduler): separate image-generation AUTO chain with UI toggle
Image models previously could not be scheduled through a priority
chain: the chat AUTO chain explicitly skips image-kind slots, and
AUTO image requests fell back to unordered registry discovery.

- config: add auto_image rules (auto_image yaml / image_rules json);
  legacy auto rules keep their meaning as the chat chain
- core: buildAutoImageChain mirrors buildAutoChain with inverted kind
  filter (image-only); SaveAutoImageRules + AutoImageRules/AutoImageChain
- scheduler: ChainImage walks the chain tier-by-tier with round-robin
  and preference ordering, skipping cooling slots
- gateway: handleImage AUTO now runs down AutoImageChain when one is
  configured (falls back to legacy discovery otherwise) and records
  the actual served model; handleAutoAPI GET returns image_rules and
  PUT accepts image_rules independently of rules
- webui: priority page gains a chat/image toggle editing two
  independent lane sets; add-slot picker filters by active kind;
  persistAuto writes only the active chain's field
2026-08-26 21:03:51 +08:00
66585549f1 fix(webui): allow image models in priority/AUTO chain editor
The priority page and key-scope pickers skipped kind=image models,
so image sources (e.g. Kwai-Kolors/Kolors) could not be placed in
the AUTO chain or granted per-key. The backend already supports it:
handleImage resolves AUTO through the chain then filters with
imageOnly, while chat requests are protected by chatOnly, so image
slots never receive chat traffic.

Image models now appear in the priority canvas, the add-slot picker
(labelled ' (image)'), and the key scope dialog.
2026-08-26 20:36:05 +08:00
e48baa1bb3 fix(webui): missing HTTP method on fetch calls with body
6 call sites used the fetch default GET while sending a request body,
which browsers reject outright ('Request with GET/HEAD method cannot
have body'). Affected flows: create key, save source, save auto rules,
upload adapter, update key model scope, and /api/chat streaming.

All now send the method their backend handlers require (POST or PUT)
with an explicit Content-Type.
2026-08-26 00:34:57 +08:00
dev
21ec8f59d8 feat: surface zero cache hits — distinguish 'missed' from 'not reported'
Live testing across the zen pool showed models report
prompt_tokens_details.cached_tokens even when the hit count is 0 (e.g.
nemotron-3-ultra-free returns cached_tokens:0, audio_tokens:0,
cache_write_tokens:0). The previous >0 guard dropped those objects, so a
cache-enabled upstream looked identical to one without cache support.

- types: PromptTokensDetails.CachedTokens always emitted (drop inner
  omitempty) so clients see cached_tokens:0 explicitly; dsh reads it as
  a 0% hit instead of 'no data'
- adapters (9): forward prompt_tokens_details whenever the upstream
  provides it (presence check instead of >0)
- Req: add cache_reported flag set when usage carried cache accounting;
  WebUI shows an amber 0% tag for reported-but-missed rows and keeps
  the em-dash only for sources that never report cache data
2026-08-25 10:04:44 +08:00
dev
24609289e8 feat: record cache hit/miss per request in audit trail and WebUI
- Req: add CacheHit and CacheMiss fields (carrying upstream cache
  accounting from either prompt_tokens_details.cached_tokens or legacy
  prompt_cache_hit_tokens)
- recordChatUsage (non-streaming): copy cache fields from resp.TokenUsage
- pumpStream (streaming): write lastUsage cache fields back onto rec at
  stream end, so streaming requests carry cache data too
- CSV export: add first_byte_ms, cache_hit_tokens, cache_miss_tokens
  columns alongside the existing latency/prompt/completion
- WebUI request-records table: add a Cache column showing hit% per row
  (green/amber tag with tooltip hit/miss breakdown; em-dash when the
  upstream reported no cache data)
2026-08-25 09:36:21 +08:00
dev
18c2385c61 feat: add per-source TTFB and tokens/s metrics to status page
- Req: add FirstByteMs field (ms to first byte, tracked for streaming)
- Stat: add FirstByteSum for aggregation
- SourceAverages(): new method computing per-source avg TTFB and tokens/s
  from the in-memory ring (300s window)
- SourceStatus: add AvgFirstByteMs and AvgTokPerS fields
- pumpStream: record FirstByteMs after first SSE chunk sent to client
- singleChat/singleChatAuto: set FirstByteMs = LatMs (non-streaming)
- handleStatusAPI: populate the new SourceStatus fields from SourceAverages()
- WebUI source table: two new columns showing TTFB (s) and Tokens/s
2026-08-25 08:24:08 +08:00
dev
334b984c25 fix(webui,gui): repair dead export button + wire chat clear; prune UI redundancy
WebUI (internal/gateway/ui):
- BUG: the export modal's custom-range button called
  downloadStatsCsvFromForm() which was never defined — clicking it threw a
  ReferenceError and nothing downloaded. Implement it: reads #exp-from /
  #exp-to date inputs and forwards to downloadStatsCsv.
- BUG-adjacent: clearChat() existed but was reachable from no control —
  add a Clear button to the chat composer so conversation reset is actually
  possible (+ cClear i18n zh/en).
- remove byte-identical duplicate html[data-theme=dark] CSS block (15 lines)
- remove 8 dead CSS rules (.keys-grid .m-model-row .scope-add/.scope-box/
  .scope-chips .scr-blocks .tag-warn .twrap) and the never-consumed
  --accent custom property
- remove 3 dead JS functions (activeTab/findSlots/scopeUncomb; lastTab decl kept)
- remove 24 dead i18n keys x zh/en (~55 lines) — legacy of the replaced
  key-scope editor, matching the removed .scope-* styles

GUI (cmd/gui/main.js):
- BUG: stopCore() set app.isQuitting=true and nothing reset it — after using
  tray 'stop core', closing the window quit the whole app instead of hiding
  to tray, and core crash auto-restart stayed disabled. isQuitting now only
  flips in restartCore (scoped) and before-quit.

Verified: go vet/test green; node --check on all three GUI js files and the
WebUI inline script.
2026-08-24 23:13:54 +08:00
690f55c6f1 feat: proactive rate limiting (RPM) + 429 short cooldown for sources
- config.go: Source add RPM field (requests-per-minute cap, 0=unlimited)
- provider.go: RecordRateLimit() — 429 uses fixed 30s cooldown, not exponential
- provider.go: Throttle() — token-bucket proactive rate limiter, spaces requests
  at 60s/RPM interval, respects context cancellation
- provider.go: ReportStatus() — 429 -> RecordRateLimit, 5xx -> RecordFailure
- provider.go: Chat/ChatStream — wire Throttle after TryAcquire
- api.go: sourcePayload + RPM, buildSource passes RPM through
- ui/index.html: add RPM input field in source editor, bilingual i18n labels
- deploy.sh: backup old binary + rollback on healthcheck failure
- provider_test.go: TestModelStateRateLimitShortCooldown, TestThrottleSpacingAndCancel
- config.yaml: sensenova rpm: 12
2026-08-24 02:00:25 +08:00
8c5279b416 fix(status): source column reflects real traffic + theme-aware tray menu
- api/status sources now carry recent_ok/recent_err (last 300s real gateway
  requests via Stats.SourceRecent) so a source actually serving traffic is
  never shown as down just because probe /models got rate-limited
- WebUI source status column repaints every 5s (no more frozen-at-first-
  render) with a manual refresh button; shows success rate + probe + cooldown
- tray menu status rows were enabled:false (GTK fixed light-grey, invisible
  on light themes) — now enabled with no-op click and nativeTheme listener
  rebuilds the menu on dark/light switches
- ignore local ops scripts (scripts/, machine-specific)
2026-08-17 17:37:04 +08:00
3061a70087 style(ui): normalize index.html formatting (expanded attrs, 2-space indent) 2026-08-16 13:15:38 +08:00
ae1e1d39b4 chore(ui): remove NapCat design-DNA references and napcat-design-dna.json; neutralize design attribution comments 2026-08-16 13:05:53 +08:00
2bc1d0e67a feat: opencode zen adapter + first-run config generation, fix stats/stream bugs
- adapters/opencode.lua: opencode.ai zen free pool adapter — sends the
  opencode client User-Agent (zen fingerprints clients by UA; non-official
  clients hit FreeUsageLimitError); pairs with api_key: public
- config: no config file ships in the repo; first run generates a default
  config at the -config path with a random admin key, loopback listen and a
  keyless zen source (config.EnsureDefault); remove config.example.yaml
- lua: seed bundled adapters from the embedded FS instead of a hardcoded
  name list
- ui: widen model kind select (chat was clipped to 'cha')
- phase 5 bugfixes: stats ms/s bucket mixing, cleanScopes nil, ctx.Err
  guards, direct-path ModelAvailable, empty stream body failure,
  bestImageModel rewrite, transform failure recording, Core.mu, timer,
  effective model for tool-calls
2026-08-13 12:25:07 +08:00
d06210204b feat: config.yaml only, OpenRouter free models, remove runtime.json config
- Move all config (auto rules, keys) from runtime.json to config.yaml
- Store now only holds runtime sources (WebUI-created)
- Add OpenRouter free models to config.yaml
- One-time migration from legacy runtime.json on startup
- Fix gateway tests for new config structure
- Update core.go with migrateFromRuntime, saveConfig, seedKeys/seedAuto
- Remove SaveKey/KeyByValue/AutoRules from Store
- Add Config.Save() with YAML marshaling
- Update WebUI admin keys visibility (show all keys including admin)
- Bump binary to 11MB with luajit
2026-08-13 10:18:19 +08:00
084c9fee2b fix: narrow-screen sidebar no longer blurred
The sidebar's backdrop-filter: blur() made the popup sidebar itself
unreadable on ≤900px screens. Disabled the blur in the responsive
media query; now sidebar has solid glass background and the
full-screen blur only applies to the backdrop behind it.
2026-08-12 16:05:52 +08:00
3a590e5053 UI重构: 采用NapCat设计语言
- 坚持单文件embed(go:embed ui/*),保留全部既有JS功能
- 从NapCat WebUI提取设计DNA并产出 napcat-design-dna.json
- 新增主色调切换:Sakura(默认)/Ocean/Violet 三套完整色板
- 全部emoji替换为内联SVG图标
- 侧边栏响应式折叠(≤900px),背景毛玻璃+动态梯度blobs
- 优化性能:降低blur强度、记录表仅渲染最近300条、支持prefers-reduced-motion、页面隐藏暂停轮询
- 修正焦点轮廓、移除方形框提示
- 面包屑+hamburger+sidebar+accent picker完整chrome
- 通过Playwright无头浏览器全链路验证:登录/导航/主题/语言/响应式/聊天/统计
2026-08-12 13:51:29 +08:00
32303e4238 fix(ui): cache the status page so switching back to home no longer rebuilds the whole dashboard
renderStatus() re-injected the entire status tab (#tab-status.innerHTML) + all charts + tables on every visit, so each tab switch back to home stuttered. Now the page is built once (pane.dataset.built) and revisits only reload live data via paintStats() + restart the 3s poll. Static parts (source table, model chips, connection box) are not rebuilt. Deployed (bak .bak.20260811y), server active.
2026-08-11 18:44:12 +08:00
788e5a3d0d fix(ui): tokens card — fix 252px height overflow, show all models with usage in legend
- .kpi fixed height 206px was too short for the tokens card (hdr + big value + two-line in/out + chart + legend) -> text overflowed; bumped to 252px with legend capped at max-height/scroll
- tokens chart: stack shows top-4 slices + lumped 'other', total uses the FULL model set (true shares, was relative to top-4); legend lists every model that has non-zero tokens (was hardcoded top-4, so small/zero slices were dropped -> 'only two models' visible)
- add kOther zh/en
- deployed (bak .bak.20260811x), server active
2026-08-11 18:39:09 +08:00
bce3f1b23d fix(ui): eliminate KPI card height jump on load — all five cards fixed at 206px, chart area flex:1
Previously the skeleton was a fixed 186px but the real cards had different natural heights (status pie 110 + legend, tokens two-line sub + legend, etc.), so replacing the skeleton caused a size jump. Now every .kpi is a fixed 206px flex column; .k-chart absorbs the remaining space (flex:1) and the canvas absolutely fills it (draw fns read container height, not a hardcoded 86/110). Skeleton .ksk is exactly 206px with the same layout, so first paint and real cards are identical in size. Deployed (bak .bak.20260811w), server active.
2026-08-11 18:02:28 +08:00
dbee5188f8 feat(ui): align skeleton height to real KPI cards + add page/tab/card entry animations
- KPI skeleton (.ksk) min-height now matches the real card (186px) with a title/value/chart bar layout so first paint doesn't jump
- global entry animations: cards fade+slide-up (fadeUp), tab panes fade+rise (tabIn), chat empty state fades in
- goTab restarts the tabIn animation on switch (reset animation + forced reflow)
- table-row animation intentionally NOT added (those re-render every 3s poll and would flicker)
- deployed (bak .bak.20260811v), server active
2026-08-11 17:55:24 +08:00
12c61317fb feat(ui): loading skeleton for the KPI row so first paint isn't a sudden blank->cards pop
The five KPI cards are built only after /api/stats returns; before that #kpi-row was empty, so the top of the dashboard appeared blank then jumped in. Now the row starts with a shimmer skeleton (5 placeholder cards with a CSS gradient animation) that paintStats replaces with the real cards on first data. Deployed (bak .bak.20260811u), server active.
2026-08-11 17:50:23 +08:00
a83abee378 fix(ui): stop rebuilding the five KPI cards every 3s poll (was causing a visible flicker)
paintStats() rewrote #kpi-row innerHTML on every 3s refresh, destroying/recreating all five cards + canvases each time -> each poll flashed. Now the card structure is built once (dataset.built guard) and subsequent polls only update the value text (data-kpi attrs) and redraw the canvases. Preserves colored success rate and two-line in/out tokens. Deployed (bak .bak.20260811t), server active.
2026-08-11 17:42:51 +08:00
5ca3b7f680 fix(ui): chat model dropdown shows source (src:model pin), tokens card stacked bar + swatch legend, tokens in/out on two lines
- chat tab model select: list each model grouped by source as 'source · model' with value 'source:model' (the unambiguous pin syntax — all 18 real models contain '-' but none contain ':'), so same-named models on different sources are distinguishable in tests; falls back to flat model list for non-admin (sources not exposed)
- tokens card: remove on-chart text, add color-swatch legend (#tb-tokens-legend) mapping bar color -> model; in/out tokens on separate lines (.k-sub.k2)
- deployed (bak .bak.20260811r), server active
2026-08-11 17:34:19 +08:00