Commit Graph

107 Commits

Author SHA1 Message Date
aa10ee8c27 fix(billing): 插件页不可达(真·空白根因)+ i18n + 溢出 + 状态页 tile
## ★ 用户报告「Billing 页还是空白」—— 上一轮的验证有漏洞
上一轮我用 goTab('billing') 直接调用验证,显示"有数据、无错误"就下了结论。
但用户是**点侧栏按钮**。真实点击路径走 goTab,而 goTab 只遍历硬编码的 TABS
常量来切换 `hidden` 类 —— 插件页不在 TABS 里,所以 #tab-billing 的 hidden
**永远不会被移除**。内容一直躺在 DOM 里(KPI/表格都填好了),只是不可见。

这个 bug 没有任何报错:注入正确、数据正确、API 200,唯一的问题是宿主页的
路由逻辑没把插件页纳入。而它是上一轮「TABS 收敛为单一常量」时留下的:
收敛让三处共用一个常量,却没让插件页进入它。

修法:goTab 同时遍历 PLUGIN_PAGES。PLUGIN_PAGES 从 const 改为 var 并**提前到
goTab 之前声明** —— const 在文件后部声明的话,goTab 的读取落在 TDZ 里,
第一次点击插件页就会抛 ReferenceError(同类问题这个文件里已是第二次)。

判据 TestPluginPagesAreReachableByGoTab 锁两件事:goTab 遍历插件页集合 +
声明在使用之前。变异验证:删掉遍历 → 红;var 改回 const → 红。

## 插件 UI 不跟随多语言
宿主的 applyI18n/data-i 只覆盖**宿主渲染的标记**;插件注入的 HTML 对它不可见,
所以整个 UI 切中文时 Billing 页还是英文。

pluginAPI 增加 lang(getter,实时值)与 onLangChange(切换回调)。
billing 页所有文案改走双语字典:KPI、表头(fresh/cache/cache%)、区块标题
(占位后由脚本填)、空态、状态页 tile 标签。切换时立即重渲染标题,
不用等下一次 fetch。

## 部分页面超出 UI 区域
#main 只有 overflow-y,插件页内容(8 列表格 min-width、长字符串)会横向撑破。
两层修:插件 pane 统一 min-width:0/max-width:100%/overflow-x:auto(第三方
任意 HTML 的兜底,与原生 pane 一致);billing 的宽表格在自身容器内滚动。

## 状态页 tile 的 TypeError(每次重绘都报)
tile 的 tick() 在 await 之后直接 getElementById(...).textContent = ...,
但状态页每次刷新都整体重建 pane,元素可能已不存在 → null 属性赋值。
await 之后重新取元素并判空。

## 顺手补的缺口
上一轮加了缓存表格列,但 KPI 卡片漏了(那次替换 assert 失败后重试只重做了
表格)—— 缓存命中率在表格里有、KPI 里没有。本次补上。

## 验证
真实浏览器(禁缓存、真实点击侧栏按钮):pane 可见、KPI 9 项、表头双语、
语言双向切换正确(Per source ⇄ 按源)、无水平溢出、无 billing 控制台错误。
生产数据:Total USD 0.566798 / 732 请求 / 降级 231 / 2.09 亿 prompt tokens。
387+ 测试全绿。

## DSL(进行中,未完)
config.BillingDSL(active + profiles + rules,rule 按 url 匹配 mode=free/
token/subscription/unpriced)与 internal/billing.Compile(url 规则 → 插件
prices 表,含峰谷窗口的形状编译 —— 之前手写 JSON 两次弄错的正是这个形状)
已落地并通过校验/编译;core 启动接线已写。profile 切换 API 与 WebUI 选择器
未做,生产 config.yaml 也尚未写 billing 段 —— 下一轮继续。
2026-10-02 11:47:24 +08:00
fbdf0dea10 fix(billing): 缓存命中统计缺失 + Billing 页空白 + 侧栏图标
三个问题都来自生产实测,不是代码审阅。

## 1. 缓存命中被计费却不被统计
网关确实从上游 usage 提取了 prompt_cache_hit_tokens(审计里能看到
cache_hit_tokens: 270104 / cache_reported: true,占 prompt 的 99.9%),
costFor() 也用它给缓存段定价了 —— 但**没有任何 bucket 记录它**。
结果:一个 99.88% 命中率的网关,报表显示 prompt_tokens 却看不出其中
多少是缓存读,也无从按源/模型/key 看命中率。

每个 bucket 现在多三个字段:
  cache_hit_tokens    命中数(按上游上报)
  cache_fresh_tokens  未命中的 prompt
  cache_reported_reqs 上游确实上报了缓存数的请求数

第三个字段是刻意的:**「零命中」与「上游根本不上报」在命中总量里完全一样**,
而它们在「缓存折扣有没有生效」这个问题上含义相反。没有它就无法区分,
只能猜。

chat.go 的 payload 之前**没有** cache_reported(审计有、插件没有),
所以任何插件侧的缓存统计都只能猜 —— 已补上。

旧 state 文件的 bucket 没有这些字段:Lua 里 nil + number 会抛错,而钩子抛错
会让**该请求完全不记账**(一个统计缺口会变成静默缺口)。add() 里做了回填。

UI 增加 fresh/cache/cache% 三列 + Cache hit rate KPI;未上报的显示 n/r 而不是 0%。

## 2. Billing 页空白:render() 引用了未定义的 s
`render(st)` 里两处 KPI 写成 `s.degraded_reqs`,ReferenceError 让整个渲染
中断,所有表格停在初始的空 innerHTML。症状是「页面加载了但什么都没有」,
而 /api/plugins/billing/state 返回 200 且有真实数据 —— 载荷完全正确,
DOM 是空的。

更糟的是 refresh() 里的 `catch (e) { /* never break the page */ }` 把错误
**静默吞掉**了:网络面板一切正常,页面什么都没有。现在 catch 会
console.error(仍然不抛,装饰性组件不该拖垮宿主页,但必须留痕)。

## 3. 侧栏图标
billing 声明 icon = "💰",而原生 tab 全是内联 SVG(stroke: currentColor)。
emoji 尺寸不对、不跟随主题。

WebUI 增加 pluginIconHTML:插件图标可以是文本,也可以是内联 SVG。
**SVG 走严格白名单**(tag + 属性都是 allowlist,不是 denylist)——
插件是在运维者浏览器里跑的第三方代码,不能"信任插件";但也不能直接拒绝
SVG,因为那是唯一能和原生 tab 视觉一致的方式。

用真实 Chromium 验证 12 个用例,全部挡住,包括 foreignObject 里嵌 HTML
命名空间 <img onerror> 这个经典绕过(整体丢弃,所以 img/onerror 也没了)。
★ node 里没有 DOMParser/jsdom,所以没法在单测里跑这个过滤器 —— 用正则近似
会得到一个"测试通过但浏览器里失效"的过滤器,这比没有测试更糟。

顺带修了过滤器的两个真缺陷:输出里嵌套了空 `<svg></svg>`,且 viewBox
是从包装元素读的(永远是 null)而不是插件自己的,所以任何自定义 viewBox
的图标都会丢失。

## 判据(新增 7 项,全部变异验证)
写「注入脚本能否正常执行」这个守卫时我错了四次:
  1. 静态扫「已声明的名字」→ 把 HTML 字符串里的 CSS 类名(class/div/td)
     全报成未定义
  2. 用 CSS 选择器解析器查样式表 → 报样式表本身坏了
  3. 只挂 process 的 uncaughtException → 脚本在 IIFE 里异步跑,错误是
     unhandledRejection,判据对原 bug 全绿
  4. 只查「有没有抛错」→ render() 开头是 `if (!st) return`,传错字段是
     **静默 no-op**:不抛、不打日志、不报错,只是页面空白
最终判据是:在 node 里用 DOM stub 真跑一遍,同时要求「无异常」且
「至少写进一个容器」,并监听 console.error。变异验证:还原 s → 红;
render 收到 undefined 字段 → 红。

表头/行列数一致性也有守卫:row() 加了缓存列而表头没加时,表格会整体错位
(cache% 落到 completion 列下)—— 渲染正常、有数据、但要仔细看才发现。

## 生产验证
重启后价目表与累计账完整保留(1.17 亿 prompt tokens)。
新请求缓存统计生效:cache_hit 947,436 / cache_fresh 888,
cache_reported_reqs 7 / 395(其余来自旧 state,正是该字段存在的意义)。
真实浏览器:表格 3 行、KPI 7 项、表头 name/cost/reqs/prompt/fresh/cache/cache%/completion、
SVG 图标 currentColor 渲染、控制台无 billing 错误。391 个测试全绿。

## 另发现一个无关 bug(未修)
首页 stats 图表抛 IndexSizeError: arc 半径为负(-2),在 ui/index.html 的
paintStats 附近。属状态页图表,不在本次范围。
2026-10-02 11:16:15 +08:00
8de1499c40 fix(ui): 插件元素注入被宿主页面重建擦除 —— 元素型注入从未真正生效
## 现象
部署示例后打开 WebUI:侧栏有 Billing 页,但**状态页上没有任何计费组件**。
插件明明声明了 elements,/api/ui-inject 也确实返回了 mount(1570 字节)。

## 根因(不是缺功能)
七个宿主页面的渲染函数都用 `pane.innerHTML = ...` **整体替换**自己的 DOM。
`renderStatus` 在 `injectPluginUI()` 之后由 refresh() 立刻调用,于是刚挂上的
plugin-el 连同整个 pane 一起被下一次赋值销毁。

时序上它**从没有过"显示一帧"的机会**:注入 → refresh("status") → innerHTML 覆盖。
所以症状是"元素从来没出现过",而不是"刷新后消失"——这正是我先前据
/api/ui-inject 返回值判定"注入正常"而漏掉的地方:**载荷到达 ≠ DOM 存活**。

Billing 页不受影响,因为它属于 PLUGIN_PAGES,走插件自有 DOM,不经宿主重建。
于是看起来像"页面注入有效、元素注入无效",把排查引向插件声明本身。

## 修法
- mountPluginElements 改成具名可重入函数,并注册进 PLUGIN_MOUNT_HOOKS
- refresh() 在**唯一出口**统一调 remountPluginElements(),而不是给七个渲染函数
  各加一次调用——后者是多一处会忘的地方,而忘记的后果是静默的
- 每个挂载点按 data-idx 幂等:宿主重绘时若该 pane 已有该元素就直接返回,
  否则插件的 <script> 会每次重绘都跑一遍,计数器静默翻倍

## ★ 验证方式换了:真实浏览器,而不是 payload
静态测试和 curl 都看不出这个 bug(载荷完全正确)。用 CDP 连本机共享浏览器实测:

  修复后:首屏 tile=1,页面重建后=1,连续重建 5 次仍=1,console 无错误
  回退后:tile 全程=0

  对照二进制(把 done 改成空函数重编译)实测首屏就是 0,
  **证明"从未显示过",不是"显示后消失"**。

## 判据与变异
TestPluginElementsSurviveHostRebuild 锁住:remountPluginElements 存在、
PLUGIN_MOUNT_HOOKS 在使用之前声明(const TDZ 会让首屏直接抛错)、refresh 挂了重挂、
挂载按 data-idx 幂等。

三个变异全部被抓住:撤掉 refresh 的重挂 / 去掉幂等守卫 / 把 const 声明移到 push 之后。
★ 第一次跑第三个变异时**判据正确地没报**,因为我的替换脚本命中了注释里的同名文本,
真正的 const 没被移动——是变异无效,不是判据有洞。换按行定位后如期变红。

## 同时补上部署示例
packaging/config.example.yaml 里补 plugin_dir 说明(之前只有 online 部署路径踩过)。
实测升级路径本身是好的:给已有配置加 plugin_dir 后,首次启动会自动 seed 内置
billing 插件,无需手工放置文件。
2026-10-02 09:04:47 +08:00
1c690611f8 feat(gui): WebUI 与 Electron 壳的插件安装/删除/禁用/编辑
## WebUI:新增「插件」页
- 列表来自 on_disk(不是 loaded 集合)——**加载失败的插件也必须显示并带错误**,
  否则一个语法错误看起来和"插件没装"完全一样
- 启用/禁用(PUT {"enabled":bool})、删除、编辑源码、安装/覆盖
- 显示 hook_errors:插件抛异常在别处毫无痕迹,没有这一栏的症状就是
  "功能就是不work"
- 插到 dropzone 与代码编辑器都做了泛型化(bindDropzone / openCodeModal),
  适配器与插件共用一份,而不是复制第二份只改 4 个 id 的函数

## TABS 收敛为单一常量
tab 清单原本是字面量散在三处:goTab、refresh()、admin-only 隐藏列表。
加一个 tab 意味着三处都要记得改,漏一处就是"路由认得但界面不显示"——
和今天早些时候 chain_step 漏报同一类静默缺口。现在只有 const TABS。

## Electron 壳:设置面板里的插件管理
渲染进程不能直连内嵌核心(没有 key、不知道端口),所以走 IPC:
  renderer → plugins:proxy → main → HTTP /api/plugins
代理是 (method, path, body) 透传而不是固定命令表:固定表每加一个端点就要扩,
而"按钮存在但什么都不做"比"没有这个按钮"更糟。透传让渲染层能调用核心将来
新增的任何 /api/plugins 路由,路径在主进程校验。

## ★ GUI 此前零测试,而本次改动就引入了三类"看起来没事"的问题
1. 引用了不存在的 CSS 类(.tag / .sm)——渲染成无样式文本
2. 引用了不存在的 helper(esc / escAttr)——那是 WebUI 的,renderer/app.js
   是独立文档,点击时 ReferenceError
3. .ghost/.primary 只在 .form .actions 作用域内生效,插件按钮在 .pl-acts 里
   于是是无样式裸按钮

补 4 个静态判据(不启动 Electron,守卫的正是"打开应用才看得见"那一类):
  TestGUICSSClassesExist          用到的类必须在样式表里定义
  TestGUIHelperFunctionsAreDefined 被调用的函数必须有定义
  TestGUIPluginPanelIsReachable  面板在 overlay 内、按钮已绑定、打开设置会加载
  TestGUIIPCPathIsConstrained    代理必须限定 /api/plugins 前缀并拒绝路径穿越

写第一个判据时我错了三次:CSS 解析器先丢最后一个 selector、再把变量块当
selector、最后漏掉复合选择器(.tb-btn.tb-close)。两次"判据自己坏了"的
教训和本项目一贯一致——**判据出错的信号是它报了一个假问题**。现在改用宽松的
token 提取 + 显式的 guiKnownUnstyled 豁免表(blob/tgl/rail 是既有无样式类,
不是本次引入,失败它们只会让判据对新工作失去意义)。

## 变异验证
  改坏唯一的 CSS 定义(.pl-empty)→ TestGUICSSClassesExist 红
  改坏 helper 名 → TestGUIHelperFunctionsAreDefined 红
★ 第一次变异我改了 .pl-broken,判据**正确地没报**——因为它还被另一条规则定义。
  这是变异选错目标,不是判据有洞;换 .pl-empty 后如期变红。

363 个测试全绿。
2026-10-02 08:47:08 +08:00
a78f7cb6c5 feat(plugin): 启用/禁用 + 磁盘列表 + 峰谷定价 + 随核心发布
## 插件管理后端
- PUT /api/plugins/{name} {"enabled":bool}   启用/禁用
- GET /api/plugins/{name}                    读源码(编辑器用,与 /state 区分)
- GET /api/plugins 的 on_disk 字段            列出目录里所有 .lua 及其加载态
- validPluginName 提取为共享函数,install/remove/read 三处共用,防止检查漂移

禁用是**运行态开关,不删文件**:插件把线上网关搞坏了、但离修好只差一行时,
运维需要把它移出请求路径而不丢失它(同 systemd mask 而非 remove 的道理)。
它**不跨重启保留**——一个悄悄比操作者意图活得更久的"禁用"本身就是个意外。

Builtin 的判定是「加载的源码与内嵌版本逐字节相同」,而不是「名字匹配」:
被改过的 billing.lua 不能被标成 builtin,否则 UI 会提供覆盖用户改动的操作。

on_disk 列表包含**加载失败**的插件。否则一个语法错误的插件在 UI 上直接消失,
运维看到的现象是"插件不见了"而不是"插件报错了"。

## 峰谷 / 时段定价
commandcode 的 DeepSeek V4 系列就是高峰 01-04 & 06-10 UTC 工作日 2 倍价
(非高峰 17h/天)。静态价目表达不了,而算错方向是**静默**的。

价目条目可带 peak = {multiplier, windows=[{days, hours}]}。命中任一窗口即乘。
★ 用 `os.date("!%H")` 取 **UTC** 小时:provider 费率表按 UTC 标注,而网关跑在
本地时区(本机 Asia/Hong_Kong)。混用本地小时会让峰谷整体偏移 8 小时,
白天算成夜间——比不做峰谷还糟。

## ★ 实现与注释不一致,被判据抓住
applyPeak 最初直接 `price.prompt = price.prompt * m`,注释写「缓存读不翻倍」。
但 costFor 里**缓存读价是从 price.prompt 派生的**,所以原地翻倍会把缓存读
也翻倍——两个折扣被叠在一起,而 provider 从没打算叠。
改成 applyPeak 只**记录**乘数,由 costFor 分段应用:fresh prompt 与 completion
翻倍,cache read 那一项不动。
只靠注释说明意图是不够的:TestBillingPeakDoesNotDoubleCacheRead 立刻红了
(0.006 vs 期望 0.003)。变异回原实现仍是红的。

## 判据(21 个计费测试全绿,新增 5 个峰谷)
  窗口恒命中 ×2 / 窗口永不命中保持静态价 / 星期不匹配不命中
  (这条正是防"用本地时区整体偏移 8 小时")/ 无 peak 规则向后兼容
  / 缓存读不随峰谷翻倍

后端部分:构建/vet/gofmt 干净,8 个包全绿。
2026-10-02 08:33:41 +08:00
d0c7465130 fix(plugin): /api/ui-inject 的 stages 漏掉 chain_step + 忽略临时构建目录
部署到线上时用隔离实例(独立端口 18099 + 独立 config/runtime/adapter 目录)
发真实请求验证,抓到的第三个 bug。

## bug:discovery 文档漏掉新 stage
handlePluginUI 的响应里 stages 是**字面写死的三个**。加 chain_step 时只改了
lua.AllStages,没改这里,于是插件作者读 GET /api/ui-inject 会看到
["request_start","routed","request_end"],**合理地得出结论:没有 chain_step 这个
stage**。stage 本身是注册好的、也确实在触发,只是没被声明。

改成从 lua.AllStages 派生——AllStages 是唯一定义顺序的地方,让它保持唯一。

★ 而我原来的测试断言 `len(view.Stages) != 3`,**断言本身是 bug 的保护伞**:
它把"三个"固化成了期望值,于是新增第四个 stage 时测试全绿、bug 上线。
现在断言改为「与 AllStages 等长且逐项相同」,并显式要求 chain_step 在其中。
新增 stage 而忘了声明,这类问题会立刻红。

## 顺带:.gitignore 补上临时构建目录
.probe/ 和 .build-work/ 是我调试时当 GOTMPDIR 和临时二进制用的,之前每轮
手工删,这轮差点提交进去 9.5MB 的二进制。

## 部署验证留档
隔离实例跑真实 chat(158 prompt / 13 completion / 128 cache_hit),计费插件
算出 0.000688,与手算 (158-128)*1e-5 + 128*1e-5*0.1 + 13*2e-5 **逐位吻合**。
★ 第一次手算我按全价算成 0.00184,一度以为插件算错了——查审计记录才看到
cache_hit_tokens。**算钱不对时先查输入再怀疑实现**,而我忘的恰是刚修的折扣。

## 验证
351 个测试全绿;变异(stages 改回硬编码三个)被 TestUIInjectServesPluginUI
抓住,3 条断言同时红。
2026-10-02 06:24:34 +08:00
cb6df0a3f0 fix(billing): 缓存命中按全价计 + 未定价流量静默记 0
部署前审计计费插件时自己找到的两个真缺陷,都会直接算错钱。

## ★ 缺陷 1:缓存命中按全价计(高估约 10 倍)
costFor 只看 prompt_tokens,不区分其中多少是缓存命中。实测(审计脚本,非推演):
1M prompt token 里 900k 是 cache_hit → **算出 10 USD**,而缓存读通常只要 1/10
价,正确值 ~1.9。agent 流量反复重放长前缀,正是缓存要让它便宜的那类流量,所以
这个偏差恰好落在最高频的流量上。

改为拆分:
    fresh  = prompt_tokens - cache_hit_tokens  → 全价
    cached = cache_hit_tokens                  → 全价 × cache_discount
cache_discount 默认 0.1(DeepSeek/Qwen/Kimi 的量级),可按条目覆盖——**折扣率是
每个 provider 的事实、不是自然常数**,所以 0.1 只是默认值而不是硬编码常量。
另外把 cache_hit 钳到 prompt 以内:适配器报出比 prompt 还大的缓存命中数时,
fresh 会变负数,凭空产生负计费 token。

## ★ 缺陷 2:未定价模型静默记 0(最危险)
没有任何价目覆盖的请求,成本记 0,而 **requests 和 token 数照常计入 total**。
于是账单看起来完全正常,只是 quietly 少报——没有任何报错,没有任何异常。
比多算危险得多:多算你会去查,少算你不会知道。

新增两个维度把这件事变成显式信号:
    unpriced_reqs    未定价请求数
    unpriced_models  按模型点名,直接告诉你价目表缺哪一行
仪表盘加一张 "Unpriced" 卡片,**这个数应该是 0**。
任何维度(source / model / key)覆盖了就算 priced。

## 修这两个时自己踩的坑
第一版把未定价统计块写在了 `local s = plugin.state` **之前十行**,
在一个全新插件上 hook 直接抛 "attempt to index global 's'",于是
**整条请求什么都没记**——计费插件能有的最坏失败方式。
是 TestBillingZeroPricesIsSafe 的 "requests = 0" 抓到的。
代价:一个计费插件静默失效,而网关日志里只有一行 hook error。

## 判据(351 个测试全绿,计费相关 16 个)
新增 5 个,全部是**具体金额**断言:
  TestBillingCacheHitsAreDiscounted          1M/900k 命中 → 1.9
  TestBillingCacheDiscountIsPerModel         覆盖为 0 / 1 两种极端
  TestBillingCacheHitClampedToPrompt         荒谬的命中数不产生负费用
  TestBillingCountsUnpricedTraffic           只数未定价的那个,且流量仍计入 total
  TestBillingAnyDimensionCountsAsPriced      源维度定价也算 priced
2026-10-02 01:14:33 +08:00
42764bc99e feat(plugin): AUTO 调度轨迹可见(chain_step stage)
被问"还有 auto 调度相关 stage 呢?"问出来的真实缺口。

## 问题
chainDrive 只返回 (resp, src, model, err),调用方只知道**最终哪个槽位赢了**。
遍历过程中算出来又丢掉的东西——哪些档被跳过、为什么跳过、哪些槽位硬失败、
哪档全忙——一律不可见。ChainErr 里其实有这些,但**只在全部失败时**才填,
而它是 error 返回值不是记录。于是:

    "tier 1 冷却所以降级到 tier 3"  ==  "tier 1 正常接单"

对插件而言 tier 只是个常量 -2("resolved by the chain"),信息量为零。而这
恰恰是优先级链存在的全部理由,也是"我那个贵模型为什么没被用"的答案。

## 做法(scheduler 侧零新依赖)
新增 TraceEvent / TraceSink,chainDrive 多一个可选 sink 参数:

  - TraceEvent 是本包的普通 struct,sink 是 func 参数 ⇒ **不新增 import**,
    scheduler 仍然可独立测试
  - sink 为 nil 时每次 emit 只多一次 nil 判断;没有插件的网关在 AUTO 热路径上
    零开销(gateway 的 chainTraceSink 直接返回 nil)
  - 事件是纯观测:scheduler 不基于它做任何分支,gateway 也不把它喂回路由/
    冷却/配额

四种 kind:tier_skip / slot_fail / tier_busy / selected,selected 每次成功
遍历恰好一次且是最后一步。顺序保证所有 step 在 routed 之前。

## 暴露给插件
新增 chain_step stage(逐个步骤),并在 request_end 载荷里加三个便于做报表的
字段:chain_walk(上限 12 步,防审计记录膨胀)、degraded、tier_served。

## ★ 计费口径(我按推荐的做,已写进文档,需要你确认)
**按实际服务的模型计费**:降级到 tier 3 仍按 tier 3 的价算,轨迹只作观测。
理由与 §7.5 的边界一致——插件只报表不执法,两套口径混在一起会引出"降级该不该
多收钱"这种无法从代码判断的争议。若要改成"按本该用的档计价",需要在 models
价目里允许按 tier 定价,这我没做,因为那是个产品决策。

## 计费插件同步消费
by_tier_served / skip_reasons / degraded_reqs 三个新维度。skip_reasons 的等待
时长做了归一(`no free slot within <wait>`),否则 busy-wait 文案一变就多一行。
降级次数在 request_end 里计而不是在 chain_step 里计:一次降级的请求要走多步,
按步计会重复计数。

## 判据(346 个测试全绿,新增 15 个)
  scheduler  6 个:正常路径只发一个 selected / 跳档+降级可见 / 硬失败与跳档
                严格区分(不可混为一谈,否则抖动上游看起来像空闲上游)/
                nil sink 安全 / 全失败时轨迹与 ChainErr 并存且不互相破坏 /
                空链不发事件
  gateway    1 个端到端:tier 1 全 500 → 插件收到 slot_fail(tier 1) +
                selected(tier 2),request_end 的 tier_served=2 且 degraded=true
  lua        2 个:降级计数与按实际模型计价 / 跳过原因归一聚合
  lua        1 个:chain_step 是真 stage 且顺序正确

3 个变异都红:去掉 slot_fail(3 个判据红)/ 去掉 tier_skip(1 个)/
去掉 degraded 字段(1 个)。
2026-10-02 01:03:39 +08:00
8c18e0c3d7 fix(plugin): request_start 在直连与生图路径上根本没触发
被"你确定功能全部正常了?你全部测试了?"问出来的。之前所有插件测试都是直接调
Plugins.Fire(),只证明 Lua 运行时没问题,**完全没验证网关有没有真的触发**——
把 handleChat 里的三处调用删掉,整个套件照样全绿,而线上一个钩子都不会跑。

补上走真实 HTTP 的端到端判据后,立刻抓到两个真 bug:

## bug 1:直连路径完全跳过 request_start
fireStart 只写在 handleChat 的 AUTO 分支里,任何指定了具体模型的请求(也就是
绝大多数请求)都不触发。修法是挪到 isAuto 判断之前,两条路径共用一次调用。

顺带修正位置语义:它在配额/模型范围闸门**之前**触发,所以插件能统计到被网关
拒绝的请求;否则插件永远只能报"被服务的请求数",算不出真实请求率。

## bug 2:生图路径三个 stage 全断
handleImage 是第三个入口,有自己的 handler 和自己的调度调用。"聊天能用"对它
毫无证明力。而生图是计费流量,计费插件看不到就等于少报。
已补 fireImageStart + 两处 fireRouted(direct -1 / AUTO -2)。
它写独立函数而不是复用 fireStart 传空 chatRequest:image 请求没有 messages
和 tools,传一个为聊天设计的零值结构会诱导后来者去读不存在的字段。

## 端到端判据(6 个,全部走真实 handler)
  TestHooksFireOnRealDirectChat      直连:三个 stage 顺序 + 真实 source/model/tokens
  TestHooksFireOnRealStreamChat      流式是另一条路径(记录由 defer 在流结束后写)
  TestHooksFireOnAutoRequest         AUTO 链:start 报 "AUTO"、routed 报**解析后**的模型
  TestHooksFireOnFailedRequest       失败请求:routed 不触发(没选到源)、
                                      request_end **必须**触发(否则计费看不到失败流量)
  TestHooksFireOnRealImageRequest    生图:type=image,第三个入口
  TestRejectedChatStillFiresRequestStart  404 拒绝也要触发 start(顺序决定的钉子)
  TestBrokenPluginDoesNotBreakForwarding   插件每 stage 都抛异常时聊天仍返回 200

## 变异验证
把 fireStart 挪回 AUTO 分支(= 重现我犯的错)→ 4 个判据红:DirectChat /
StreamChat / FailedRequest / RejectedChat。恢复后 336 个测试全绿。

这两个 bug 都属于"读代码看不出来"的类型:fireStart 那一行就在 handleChat 里,
看着挺像那么回事,只有真的发一个请求才知道它没被调到。
2026-10-02 00:49:32 +08:00
a51a6811a6 feat(plugin): Lua 插件机制 + 计费插件 + 插件文档
插件 = plugin_dir 下的单个 .lua 文件,做两件事:挂请求流水线的钩子、在启动时
贡献 WebUI 界面(整页或往现有页面追加组件)。两者独立。

## 流水线 stage(三个)
  request_start  已解析鉴权、未选源
  routed         已选定 (source, model)、未发往上游
  request_end    每请求恰好一次,带最终计量
request_end 挂在 gateway.writeRec——四条入口路径(直连/AUTO × 流式/非流式)的
唯一汇合点:既不漏(流式 token 只有流结束才知道)也不重。

## 计费插件(plugins/billing.lua,默认 seed,开箱可用)
源 / 模型 / 密钥三个维度定价。token 价优先级 keys > models > default;per_request
固定价是**叠加**的(生图模型可以既算 token 又收固定费)。单位是 USD/单 token,
即各家 provider 的公布口径。累计 total / by_source / by_model / by_key / by_day。
失败请求保留 token 费用、丢弃固定费(可经 count_failures 翻转)。
界面 = 一个独立页 + 状态页顶部一块总开销 tile。

## 一个明确的设计边界
计费插件**只报表,不执法**。网关自己的配额会计(stats.go,入口强制)才是限额
权威,插件不参与任何路由/配额决策。两套独立会计若对不上,比一套功能略少的
更糟。

## ★ 中途改掉的一个根本设计错误
最初让插件复用适配器的**弹性 worker 池**(多状态)。这对适配器是对的(它们无
状态),对插件是错的:计费插件往 plugin.state 累加,多状态意味着总量被劈成
几份;而 SetState 写价格只写进其中一个 worker,钩子恰好跑到另一个时**所有请求
按 0 计费**。改为**单状态 + 互斥锁**。代价写进文档:钩子必须短、同步、不阻塞,
卡住的钩子会卡住所有插件的钩子。
这个 bug 是测试逼出来的——先写了 SetState+Fire 的用例,数字全是 0 才挖出来。

另一个连带缺陷:只带 prices 的 PUT 会整体替换 state,把累计量清零。改为
prices/state 分离——prices 是配置、state 是历史,改价不动账。

## 撞到的三个 Lua 绑定的坑(都写进注释)
  - SetGlobal **会 pop 栈**:连着调两次,第二次从空栈取,赋成 nil
  - GetField 索引越界是 **SIGABRT 整个进程**,不是 panic,recover 救不了
  - Call(nargs, n) **不接受函数索引**,它调的是 nargs 个参数正下方那个;
    传索引会调到参数上("attempt to call a table value")
另外 GetField/SetField 用绝对索引,SetTop(0) 之后必须重取。

## 错误隔离
钩子 error() 不影响转发:捕获 → 记进 hook_errors → 跳下一个插件。适配器出错
会让源进冷却,插件出错**零惩罚**——插件是可选功能。/api/plugins 的 hook_errors
让"坏掉的插件"可见而不是静默消失。

## 界面注入
GET /api/ui-inject 一次返回所有插件的扩展(侧栏需要全部 page 才能建好)。
WebUI 在首次 render **之前** await 注入:先插 HTML 再重建 <script> 让它执行
(innerHTML/template 插入的 script 不会执行,这正是要的效果——避免脚本跑在
自己 DOM 之前)。注入失败不影响仪表盘。
browser 侧 pluginAPI 暴露 fetchState / postState / onTabShown。

## 文档
docs/plugins.md —— 快速上手、加载与热更新、三个 stage 的完整字段表、界面扩展、
状态与 HTTP API、运行时约束(单状态/异常隔离/内置函数)、计费插件的定价与
计费策略、排错表、与适配器的对比表。

## 判据(328 个测试全绿,插件相关 33 个)
  - 计费断言的是**具体金额**(0.00625 / 0.0402 / 0.0075…),不是"能加载"
  - 4 个变异都红:钩子异常不隔离 / prices 清空累计 / 忽略 key 优先级 /
    毫秒时间戳不换算
  - UI 侧 6 个判据把注入顺序、script 执行时机、pluginAPI 名称、tab 路由、
    anchor 四种形式、失败非致命全钉住
  - 鉴权:state 读任意角色、写仅 admin
2026-10-02 00:37:29 +08:00
a2e1adc2d8 fix(gemini): endpoint 自相矛盾导致预置模板必失败
gemini.lua 里 adapter.endpoint = "/v1/models",而它自己的注释写的是
  POST /v1/models/{model}:generateContent
两者矛盾,而 Go 侧是静态拼接(provider.URL = base_url + endpoint),拼不出
模型名。预置模板 "Google Gemini"(base_url=.../v1beta)于是会 POST 到
  https://generativelanguage.googleapis.com/v1beta/v1/models
既多一段 /v1,又缺 :generateContent——那是 Gemini 的模型**列表**端点,
对 POST 返 405。所以任何用户从模板建这个源,拿到的都是必定失败的源。

实测确认影响范围:线上 21 个源里没有 gemini(openai×15 / trae / sensenova /
opencodezen / deepseek / anthropic / agentrouter),所以是潜伏缺陷。

修法:endpoint 改成模板 `/v1beta/models/{model}:generateContent`,新增
provider.ChatURL(model, stream):
  - 用 **PathEscape** 替换 {model}——模型 id 进的是 URL 路径,不转义的话
    一个 "/" 就会静默指向另一个资源(判据里用 RequestURI 而非 URL.Path
    断言,因为后者是解码后的,看不出 %2F);
  - 流式把 ":generateContent" 换成 ":streamGenerateContent"(同一个路径、
    不同动词,也在路径里)。替换刻意只认这个精确后缀,免得别的适配器
    仅仅提到这个词就被改写;
  - source 自己设的 endpoint: 仍然优先,模板被整体跳过。
Chat / ChatStream / probeChat 三处调用点改为传本次请求真实的 model——AUTO
按槽位把 req.Model 钉死,所以 URL 必须跟随**请求**的模型,用源默认模型会让
多模型源每次都打同一个(还记到别的模型的账上)。

判定静态 endpoint 的其他 10 个适配器零影响(TestNonGeminiEndpointsAreUntouched)。

顺带:Stats 的 mutex 不是可重入的,导出方法自己加锁、*Locked 后缀要求调用
方持锁。持锁调导出方法会死锁——我的探针真卡死过一次(直到 10 分钟超时)。
补上 LOCKING 注释,并加判据把这条规则钉住(含一个 20 秒上限的行为判据,
让未来的重构撞死锁时快速失败而不是拖满整个套件)。
2026-10-01 23:37:17 +08:00
18e422a943 fix(sources): 源编辑不再清零 proxy_url / api_key_env / timeout
bf0657b 修好了 api_key,但 upsert 仍会重写整个源,于是请求无法表达的字段
一律被重置为零值。这四个字段的后果都不是"少个配置项":

  - api_key_env 丢失 ⇒ 盘上无明文密钥的源变成无凭据源,写操作返回 200,
    下一次调用上游才 401。而 README 恰恰把这个特性当作卖点在宣传。
  - proxy_url 丢失 ⇒ 一个走代理的上游变成直连(或反之),且完全无声。
  - timeout / queue_timeout 丢失 ⇒ 退回默认 120s / 60s。

触发路径不是只有脚本:WebUI 的 saveSource() 发的 payload 只含表单上的
11 个字段,而 editSource() 表单里根本没有这 4 项 ⇒ 运维在界面上改个并发数
就会静默清掉它们。

修法用「存在性」语义而不是「空即继承」:

  - 不传   → 保留已存值(部分更新的客户端要的就是这个)
  - 传了   → 覆盖,包括传空串表示清空

api_key 刻意保留它原有的「空即继承」规则,不跟着改成指针:该规则已随
v1.7.6 发布,脚本依赖它;而凭据丢失比代理丢失严重得多。两个字段的失败
模式相反,所以规则相反——这一条写进了两处注释。

另一处是差一点的:config.Source 把 Timeout/QueueTimeout 标成 json:"-",
所以 reveal 接口的结构体序列化**根本不返回它们**。表单读不到 → 输入框
恒空 → 而输入框每次都回传 → 每存一次就把 timeout 清零。等于把刚修好的
丢字段换个方向又造了一个。因此 reveal 分支现在显式返回 duration 字符串。

判据:
  - TestWebUIEditPayloadPreservesRoutingFields 用的是 WebUI 真实 payload
    的逐字节副本,并同时断言"确实改动的字段生效",否则"什么都不写"也能过
  - TestSourceEditPreservesAPIKeyEnv 单独盯 api_key_env(唯一造成凭据丢失的)
  - CanBeSet / CanBeCleared 分别盯两个方向:只有"空即继承"的实现过不了
    CanBeCleared(清空代理框会永远保留旧代理)
  - 持久化判据**重新加载 config.yaml 并按语义比对**:300s 会被重新序列化成
    5m0s,按字符串匹配是假红(我自己先踩了一次)
  - TestSourcePayloadCoversEveryEditableField 用反射卡住"这一类":新增
    Source 字段而没接到 API 上时立刻变红。反射查结构体而非 marshal 结果,
    因为指针 + omitempty 会合法地从序列化输出里消失,那正是"未提及"信号
  - 三个 UI 契约判据把 JS 侧也钉住(表单必须回传、必须从 reveal 读)

变异验证(每次都先确认 build 通过,再数红格):
  1. 去掉覆盖逻辑        → 7 个判据红
  2. 改成"空即继承"      → CanBeCleared + ClearIsScoped 红
  3. reveal 不返回 duration → TestSourceRevealExposesDurations 红
  4. 表单不回传 api_key_env → TestUIEditFormRoundTrips... 红

顺带修正 /api/v1 索引:DELETE /api/keys 的路径段写的是 {name},实际是 key
本身;PUT /api/keys/{key} 实现了却没列。

全量 + vet + race 全绿;WebUI 内联脚本过 node --check。
2026-10-01 20:36:48 +08:00
bf0657bb84 fix(sources): implement PUT and stop partial edits from clobbering api_key
Two defects on the admin source write path, both found while adding a model
to a live source by hand.

PUT /api/sources/{name} was advertised in the API index but never
implemented — handleSourcesAPI only switched on GET/POST/DELETE, so the
documented update verb answered 405 while the POST upsert behind it worked.

POST is an upsert that replaces the whole source, so a partial edit that did
not carry api_key persisted an empty or placeholder credential. The source
kept its name, base_url and models, the write returned 200, and the source
then answered 401 on the next request — long after the writing script exited
0. The WebUI had been routing around this by loading the real key through
?reveal=credentials; any script or partial update went straight into it.

- implement PUT, taking the name from the path and rejecting a body name
  that disagrees rather than silently resolving to one of them
- inherit the stored credential when api_key is omitted or sent as the
  literal "__KEEP__"; an explicit new key still rotates
- an empty api_key on a source that does not exist yet stays empty, since
  credential-less local upstreams are legitimate
- add model_ids, an additive shorthand, so "add these models" never has to
  read and echo the existing list back
- align the API index with the implementation

The model_ids merge had a first cut that dropped the existing list when the
request carried no models field; TestSourceModelIDsIsAdditive caught it.

Verified by mutation: removing PUT turns three tests red, flattening
resolveAPIKey into a pass-through turns TestSourceUpsertKeepsAPIKey red
on both subtests, and making model_ids replace instead of merge turns
TestSourceModelIDsIsAdditive red.
2026-10-01 18:15:29 +08:00
70f1c879bd fix(startup): 密钥警告改读真实生效的 key 集合
启动时那条「gateway_keys is EMPTY — without a key every request is rejected」
读的是 legacy 的 cfg.GatewayKeys 段,而鉴权实际用 cfg.Keys(core.ListKeys)。
seedKeys 首次启动把 gateway_keys 搬进 keys[] 之后,YAML 里那个列表就不再
被鉴权使用。于是在它被清空(例如轮换掉 starter key 之后)而 keys[] 仍有
7 把可用 key(含 admin)时,进程每次启动都谎报「所有请求都会被拒绝」。

实测:生产日志出现该警告,而同一个 key 请求 /v1/models 返回 200。

- main.go 改为检查 c.ListKeys(),文案改成不绑定字段名。
- 顺带删掉 gateway.New 的 gatewayKeys 参数:函数体从未使用它,
  只读 ListKeys(),留着会继续诱导人以为鉴权来自那个列表。

判据:e2e/TestStartupWarningReflectsRealKeysNotLegacyList —— 构造
「gateway_keys 空 + keys[] 有 key」的真实形态,先断言该 key 确实能鉴权,
再断言日志里不再出现那句谎报。变异验证:回退成 GatewayKeys() 即变红。
2026-09-28 23:42:48 +08:00
0121d23f91 fix(tokens): 流式统计改用上游真实 usage,图片不再记 token
两处 token 单位错误,均影响 per-model 配额计费:

1. 流式路径的 prompt/completion 只是「字节÷3」估算。
   pumpStream 明明收到了上游最后一帧的真实 usage,却只发给客户端、
   从不回写审计记录,于是配额按估算值扣。生产实测同一请求:
   上游 prompt=37/completion=179 → 记账 27/262,prompt 低估 1.4x、
   completion 高估 1.5x(双向失真)。同模型流式 completion 中位数
   是非流式的 4-27 倍。非流式路径本就用真实值,两路不一致。
   修法:lastUsage 非零时写回 rec.Prompt/rec.Compl,估算降为兜底
   (上游不报 usage 时仍保留原估算行为)。

2. 图片请求把「图片张数」记成 completion_tokens。
   rec.Compl = int64(len(resp.ImageData)),len 是切片长度即张数
   (生产 38 条 image 记录全是 1),且被计入 token 总量。
   图片生成无 token 概念 ⇒ 新增 Req.ImageCount 独立字段,
   Prompt/Compl 归 0;UI 记录表 image 行改显示张数(新增 i18n thImgs)。

顺带补 TestUILocaleKeyParity:此前无人校验 zh/en 键集合一致,
单边加键不会报错,只会显示原始键名。

新增 token_units_test.go(定值上游 6 项),做过变异验证:
回退修复实测复现 stream=16/173 vs chat=44/100、image completion=3。
2026-09-28 22:19:12 +08:00
c51066f0b6 refactor(quota): 配额改为按模型,删除整钥总配额
用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
2026-09-27 19:02:13 +08:00
652842783f test(gateway): 补配额桶的真实并发竞态判据
per-key 配额桶是共享 map:每个被记录的请求写它,每个配额检查读它。
单线程单测完全看不到这里的竞态,只有让多个 goroutine 同时读写才有效。

key_quota_concurrency_test.go:64 goroutine 跑 2 秒,并发 Record +
KeyWindowTokens + KeyWindowReqs + KeyWindowModelTokens,其中一条路径
在中途 opt in 惰性创建的 pinned 桶(那条路径一次改两个桶 map)。

go test -race 结果:零 DATA RACE,5444 万 token 全部入账。
全仓 -race(./...)亦全绿。

(cherry picked from commit 18cfd6d32b)
2026-09-27 18:44:41 +08:00
a21ae84cbe perf(gateway): 拒绝路径只判定一次 + 补配额交互判据
复查后修掉一个自己引入的缺陷,并补上此前缺失的交叉场景验证。

## 修复:拒绝路径重复判定

4 个入口原本先 checkModelScope(判是否为空)再 writeScopeReject
(内部又 checkQuota 一次)。即每个【被拒】的请求要跑两遍配额统计,
且两次之间用量可能变化 —— 判定与响应存在理论竞态。

改为 checkQuota 一次判定直接把 *quotaRejection 交给 writeReject,
消息与 Retry-After 都来自同一次读,不再有二次求值。
checkModelScope 保留(只需知道放行与否的调用方仍可用)。

## 补判据:此前完全没验证过的交叉场景

1. TestKeyQuotaWinsOverSlotQuota —— key 配额与 AUTO 槽位配额是两种
   不同作用域的限额(槽位是全网关共享的上游预算,key 配额属于单个
   调用方)。两者同时耗尽时必须报【key 配额】:报槽位配额会被表述成
   「无可用容量」,读起来像上游故障,而调用方能处理的恰恰是 key 配额。
2. TestUncappedKeyNeverBlockedByEmptyScope —— 只配模型范围、不配配额的
   key(生产上 5 把 user key 全是这种)绝不能被槽位检查误伤。

## 复查补测的实测数据

配额检查的真实开销(每请求一次,走完整 checkQuota 路径):
  配了配额    149 ns  0 allocs
  未配配额     42.6 ns 0 allocs   <- 生产上 5/7 把 key 是这种
  admin key    37 ns  0 allocs

未配配额的 key 只付 FindKey 的开销、根本不碰桶。相对一次 LLM 请求
(秒级)可忽略。

生产配置副本(7 key / 16 源 / 真加密凭据 / 真上游)实测:
- 100 并发 -> 50 成功 / 50 容量拒绝,RSS 19.9 -> 25.8 MB
- 生产形态桶内存(7 key x 8 model x 2 源 x 40 天满 retention)
  = 3.73 MB,占 ~32MB 预算的 11%
- 配额记账与 stats 一致:配 63000 配额后报 64062/63000
- **跨重启存活**:重启后从审计日志回放,仍报 64062/63000 并拦截;
  未配配额的 key 仍 200

(cherry picked from commit 9811654b3e)
2026-09-27 18:44:41 +08:00
d072a03c9a perf(gateway): 配额桶扫描改为窗口化 + pinned 桶惰性创建
审查本特性线的性能时发现两个问题,均有实测数据。

## 1. 窗口查询是全扫,代价落在每个请求上

sumBuckets 原来遍历整个 map(最多 960 个小时桶),实测 5.9us/op。
配额检查在每个请求上跑 2-3 次(key 总 token、key 请求数、scope token),
于是单请求多付约 18us。

注意这**不是本改动引入的成本**:main 上既有的 WindowTokens 同样是
5907ns/op(全扫)。是本改动让它在请求路径上被调用得更多。

改为只遍历窗口可能覆盖的桶(键是整点小时,范围是精确的,不是采样):
- 24h 窗口 200ns -> 55ns
- 1h  窗口  80ns -> 49ns
- 30d 窗口 5.9us -> 3.9us(720 次查找,只有配 month 配额时才走到)

等价性由 TestSumBucketsMatchesFullScan 保证(400 组随机桶位置 x 6 种
窗口,对全扫逐项比对)。★ 第一次写错成 floor,判据立刻抓到:
30 天窗口报 8878 而全扫是 8649 —— 正确是 ceil。

## 2. pinned 桶无条件创建,内存最坏 26.7MB

每条记录写两个桶:裸 model 与 "source::model"。但 pinned 桶只有
「配额里显式 pin 了 source」时才会被查。

实测最坏情况(20 key x 8 model x 3 source x 40 天全 retention):
HeapAlloc 26.67MB —— 而 README 宣传「16 源生产实例 ~32-35MB」,
等于吃掉 80% 内存预算。

改为惰性:只有 KeyWindowModelTokens 带 source 查询过某个 (key, model)
之后,才开始维护它的 pinned 桶。

  20key x 8model x 3src   26.67MB -> 8.87MB  (-67%)
  5key x 6model(真实)     3.36MB -> 2.26MB  (-33%)
  5key x 12model            5.82MB -> 3.59MB  (-38%)

代价:配 pinned 配额之前发生的用量无法事后按源拆分(记录里虽然有
Source,但桶只存了裸 model),所以 pinned 配额的首个窗口可能少算。
已在代码注释与判据中写明。

## 其余实测

  Record        main 基线 275ns/429B/3allocs -> 283ns/429B/3allocs
                (+8ns,分配数不变;3 allocs 来自 ring buffer)
  配额检查全路径  115ns / 0 allocs(每请求新增)
  纯读路径      6.5ns / 0 allocs

## 100 并发调度/拒绝压测(真实进程 + 可报并发峰值的假上游)

  容量 100(4+96),0.15s/请求,100 并发  ok=100 fail=0   上游峰值 42
  容量 100,0.15s/请求,200 并发          ok=200 fail=0   上游峰值 97
  容量 8,3s/请求,100 并发               ok=8   fail=92  上游峰值 8
  容量 8,0.6s/请求,100 并发             ok=32  fail=68  上游峰值 8
  容量 8,0.6s/请求,40 并发              ok=32  fail=8   上游峰值 8

上游峰值恒定不超过 max_concurrent,容量拒绝返回 503 + busyWait 有界
等待(约 2.6s)。main 基线在同条件下 ok=32 fail=68、上游峰值 8、
延迟分布相同 —— 配额改动没有触碰调度/拒绝路径。

配额拒绝单独验证(低并发避开容量拒绝):req_quota=50 用尽后
100 并发全部 429 rate_limit_exceeded + Retry-After: 1661,
**延迟仅 19-21ms**、上游 total 未增加 —— 配额在入口廉价拒绝,
不占用任何上游槽位,与容量不足的昂贵等待形成明确分工。

## 判据

key_quota_perf_test.go:4 个基准 + 2 个判据(sumBuckets 等价性、
pinned 桶惰性)。3/3 变异全被抓(无条件建 pinned 桶、firstHour 用
floor、keyHour 不再写)。

(cherry picked from commit be11a06a46)
2026-09-27 18:44:41 +08:00
f814ff7468 fix(webui): 修 7 处弹窗关闭错对象 + 模板管理器变量遮蔽
上一提交只修了自己新加的两处弹窗,全站其余 7 处是同一缺陷:所有对话框
共用 id="modal-wrap"(CSS `#modal-wrap:not(:empty){display:flex}`)且可以叠
加(seed-key 提示就盖在密钥页上),而
`const w = $("#modal-wrap"); w.remove()` 移除的是**文档里第一个**,不是用户
刚提交的那一个。

逐处改为两种安全写法:
- 能拿到按钮的(saveSource / saveTemplate / sortScopeSave / scopeSave /
  keyQuotaSave / createKey / downloadStatsCsv / downloadKeysCsv /
  scrAddFromForm):`btn.closest("#modal-wrap")`
- 拿不到按钮的:新增 `closeTopModal()` 取**最后一个**(用户看到的那个),
  并作为所有 `if (w) w.remove()` 之后的兜底
- 顺带把 sortScopeSave / downloadStatsCsv / downloadKeysCsv / scrAddFromForm
  的签名补上 btn / this 参数 —— 否则 .closest 恒为 null,表单永远不关

同时修一个相邻的既有 bug:`openTemplateModal` 的 `.map((t) => ...)` 用 t 做
循环变量,模板字面量里又调 t("srcEdit"),t 被遮蔽成对象 ⇒ 打开模板管理器
直接抛 `t is not a function`,整个弹窗渲染失败(main 上就有,git show 确认)。
参数改名 tpl。修后模板管理器完整渲染(浏览器实测:DeepSeek / 智谱 / Kimi /
SiliconFlow 各行 + Edit/Delete 按钮文案全部正常,零异常)。

判据从 2 处扩到全量(internal/gateway/ui_quota_contract_test.go,+3 例):
- 全文档扫描:任何 `.remove()` 配裸 `$("#modal-wrap")` 即失败
- 9 个关闭对话框的处理器必须走 .closest 或 closeTopModal
- closeTopModal 必须取 all.length - 1(首尾颠倒就是原 bug)
- 用 .closest 的处理器,其签名必须真的有 btn 参数 —— 否则查找恒为 null,
  表单永远不关
5/5 变异全被抓:sortScopeSave 去 btn 参数、closeTopModal 取第一个、scopeSave
退回裸选择器、saveSource 退回裸选择器、downloadKeysCsv 删兜底。

浏览器实测(共享 Chromium CDP,真实进程,每处都插入一个「decoy」弹窗
占据文档首位,复现原 bug 的触发条件):
- scopeSave / keyQuotaSave / scrAddFromForm / saveSource / saveTemplate
  五个处理器:自己的表单关、decoy 保留 ✓
- 零 JS 异常
★ 测试自身踩了两个坑,都不是代码问题:① `document.querySelector(sel) && .click()`
  在 CDP 里求值为 undefined,改成箭头函数;② 编辑模板时没填名字就点保存,
  saveTemplate 因 `if (!nm)` 早退、fetch 零调用 —— 一开始我把这个误读成
  「修复失效」,加 fetch 拦截 + 读 #s-name 的值才定位到是测试数据缺失。
  ⇒ 「点按钮没反应」要先分清是「事件没触发」「请求失败」还是「早退」。

(cherry picked from commit b5c3fea0bb)
2026-09-27 18:44:41 +08:00
ef631b43dd feat(webui): 密钥配额表单 + 修复弹窗关闭错对象
功能:让 per-key 配额在 WebUI 里可配置可见,之前的实现只有 API 与
config.yaml 能配。

- 密钥卡片头部显示配额徽标(token / 请求数 + 重置窗口),admin key
  不显示编辑入口(服务端本就永不受限,给入口只会让人以为配了会生效)。
- 新增「配额」编辑弹窗:token 配额、请求数配额、重置周期(复用既有
  的 period 词表与 n-hour 联动),预填从 canvas 的 data-* 读。
- 创建密钥弹窗同步加配额字段;选 admin 角色时自动禁用(同样因为服务端
  忽略 admin 的配额)。
- 「我的密钥」页新增 KEY-WIDE QUOTA 列,用户能看到自己这把 key 的预算。

修一个真 bug:保存弹窗用 $("#modal-wrap") 关闭自己,而全站弹窗共用这个
id、且可以叠加(seed key 提示就盖在密钥页上)。实测(共享 Chromium
CDP,seed 提示与配额弹窗共存)确认:保存后被移除的是 seed 提示,配额表单
反而留在屏幕上 —— 症状是「保存了但弹窗没关」,指向的方向完全错。改为用
点击的按钮 btn.closest("#modal-wrap") 解析自己的弹窗。createKey 有同样
问题,一并修。既有文件里另有 7 处同样写法,未动(不在本次范围,且新判据
只对本次改的两处断言,避免误伤)。

判据新增 internal/gateway/ui_quota_contract_test.go(6 例):
- 两个表单必须用 .closest 解析自己的弹窗
- putScope 必须带上 4 个配额字段(API 视其为指针,省略=清空预算)
- 创建请求必须真的发出配额字段
- **数据流判据**:徽标要真读 k.token_quota 等、编辑表单要真读
  canvas 写的 data-kquota 等。只查字面量存在会漏 —— 字段躺在死分支里
  判据照样通过(这是本轮实际踩到的:keyCapBadges 经 keyPeriodSuffix
  间接读 k.period,被判据抓到后我把读取显式化而不是放宽判据)
- 弹窗扫描先剥注释,否则修复说明里引用的字面量会被当成违规
- 复用既有 ui_contract_test.go 的 jsFunctionBody(大括号配平);
  自己第一版用 2000 字符固定窗口,被长注释顶开后仍在窗口外命中后面
  函数的同名字段,读起来像通过 —— 窗口法在这里是假判据

7 个变异全部被抓(unsafe 关闭、putScope 丢字段、createKey 丢字段、
徽标不读字段、canvas 不写 data-*、kq-hours 改名、周期词表缺项)。

浏览器实测(共享 Chromium CDP,真实进程 + 加密配置):
- 徽标渲染 1.0K·1h / 5×·1h;编辑框预填 1000/5/hour,hours 框按周期联动
- 保存后回读 250000/77/nhour/6,徽标更新为 250.0K·6h,toast Saved
- 零 JS 异常
- **关键回归**:编辑模型 scope 后配额仍是 250000/77/nhour,未被清空
- 创建带配额的 key,服务端确认 {t:50000,r:300,p:week,role:user}
- user 视角「我的密钥」显示 777·1h 与 9×·1h

文档:README.md / README_EN.md 补「密钥用量配额」小节(配置示例、
周期词表、429 语义、admin 豁免、整点分桶最晚晚 1 小时释放、PUT 的
省略 vs 0 语义、429 响应样例),特性列表各加一条。

(cherry picked from commit ce66c7f6c2)
2026-09-27 18:44:41 +08:00
9c3aabb7f9 feat(gateway): per-key 用量配额(token + 请求数)与重置周期
问题:密钥控制只能限制模型范围。实测发现三个缺陷,其中前两个让
per-model token_quota 在真实链路上从未生效:

1. 桶键不含 key。scopeTokens 调 WindowTokens(model, source, win),
   桶键是 model / source::model,与调用方无关。实测两把 key 各用
   1000 token,窗口报 2000 —— A key 的额度被 B key 消耗。
2. 无 source pin 的桶永远是空的。真实记录 Source 总被填上,桶键存成
   "deepseek::m1",而无 pin 的查询找 "m1" —— 读到 0,永远 < quota,
   配额形同虚设。实测 WindowTokens("m1","",1h)=0 而 pinned=2000。
3. AUTO scope 走 KeyTokens(key),是全时段累计、永不重置。实测 30 天
   前的 200 token 仍计入 1 小时配额(报 210 而非 10)。配了
   period: hour 也不会每小时归零。

生产 5 把 user key 全是 token_quota: 0,所以前两条一直没暴露。

改动:
- Stats 新增 per-key 小时桶 keyModelHour(key → model → hour)与
  keyHour(key 总量)、keyReqHour(请求数),retention 40 天,与既有
  modelHour 对齐以覆盖最长的 month 窗口;LoadAudit 走 aggregateLocked,
  所以窗口用量跨重启存活。modelHour 保持 key-blind:它服务的是 AUTO
  槽位配额(限制整个网关对某槽位的消耗),语义不同,不应被 per-key
  改造污染。
- 每个请求写两份模型桶:裸 model 与 source::model。无 pin 的 scope
  条目读前者,有 pin 的读后者。
- GWKey 新增 TokenQuota / ReqQuota / Period / Hours:整钥配额,
  跨该 key 所有模型共享一份预算;ReqQuota 覆盖持续请求量(源上的
  RPM 只管突发)。
- 配额耗尽返回 429 + Retry-After(rate_limit_exceeded),而不是 403:
  403 让客户端以为这把 key 永远不能用该模型,直接放弃;429 + 等待
  才能在窗口重置后自动恢复。模型越权仍是 403。
- admin key 永不受配额限制 —— 否则操作者会把自己锁在门外。
- 周期词表在写入时校验,拼错的 period 被拒绝而不是静默当成永不过期
  (那与操作者输入的意图正好相反)。
- PUT /api/keys 的配额字段是指针:省略=保留原值,显式 0=解除限制。
  否则只改模型范围就会悄悄清空预算。

判据 3 个文件 24 例,9 个变异全部被抓:key 隔离、pin 桶缺失、
AUTO 周期、key-blind 退化、429→403、admin 被限、PUT 清空配额、
Validate 失效、pinned 桶缺失。前三个变异最初漏网 —— 判据只测了
Stats 层没测接线,补了走真实 HTTP 的接线层与 API 层判据后抓住。
端到端验证:真实进程 + 加密配置往返,配额字段与 enc:v1 密钥均正常。

(cherry picked from commit 5306251840)
2026-09-27 18:44:41 +08:00
f6baa13583 feat: WebUI 改为依赖 /api/v1,UI 与 agent 共用一套 API 契约
- sources / sort / keys 三个页面的数据源从 /api/sources 切到 /api/v1/sources
  (写操作仍走 /api/sources:v1 是只读门面,不做变更)
- 编辑弹窗改用 /api/v1/sources/{name}?reveal=credentials(admin-only)取明文 key。
  这是必须的:表单要整体回传源,若不回填 key,改个端口就会把 key 清空。
- 遮蔽视图仍是默认,只有显式 reveal 才返回明文

端到端验证(真浏览器 + 临时实例,非仅 API 测试):
- sources/sort/keys 三页实际发出 GET /api/v1/sources,0 console error
- editSource('demo') → reveal=credentials,#s-key 与 #s-url 正确回填
- 写入往返:改 base_url /v1→/v2 后重开,key 仍在(未被清空)
- 落盘 api_key 明文残留 0、密文 1

测试:+1(reveal 必须 admin,否则任意 user key 可读全部凭据)
变异验证:reveal 去掉 admin 校验 → 403 断言变红

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 926b9f6565)
2026-09-27 17:12:03 +08:00
fd03ef6e4a feat: 密钥静态加密 + /api/v1 agent 管理 API
密钥加密(写侧封存 / 读侧解封)
- config.yaml 的 sources[].api_key、sources[].headers、keys[].key 落盘即
  AES-256-GCM 密文(enc:v1: 前缀),master.key 复用 runtime store 那把
- 内存里永远是明文:鉴权比对、API 返回新建 key、WebUI 编辑回填都不受影响
- 启动时一次性封存现存明文(幂等,已封存则不写盘);-check 不写文件
- UpsertSourceInYAML 增加 box 参数,新加的源不再以明文落盘
- 解密失败改为硬错误:原先 MustDecrypt 返回密文会被下次 Save 二次封存
  (实测:源 key 18→20、静默损坏),现在启动即失败且配置分毫不动

/api/v1:面向 agent 的管理 API(WebUI 零影响)
- GET /api/v1            机器可读索引,列出每个端点的方法/权限/用途
- GET /api/v1/overview   一次调用看全貌:源 + AUTO 链 + 密钥数 + 健康度
- GET /api/v1/health     仅健康快照
- GET /api/v1/models     按源分组的可路由模型清单
- GET /api/v1/sources[/{name}]  凭据遮蔽后的源
- GET /api/v1/auto       调度链与实时槽位状态
- GET /api/v1/keys       admin only,密钥元数据,绝不回显密钥本身
- 沿用同一套网关 key 鉴权;读端点任意角色,写仍需 admin

测试:15 个新用例(含负向:泄密、越权、写操作必须被拒)
变异验证:maskKey 不遮蔽→红、去掉 admin 校验→红

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit ad28a924a5)
2026-09-27 17:12:03 +08:00
97bb9c6ef4 feat(opencode): 透传 completion_tokens_details.reasoning_tokens 与上游 cost
回答「opencodego 的用量与费用透传呢」时逐字段核对上游产出,发现 usage 漏了
一项、费用整项丢失。

## 上游实际发什么(实测 opencode.ai/zen/go/v1)

  {
    "choices": [...],
    "usage": { "prompt_tokens": 37, "completion_tokens": 40, "total_tokens": 77,
               "prompt_cache_hit_tokens": 0, "prompt_cache_miss_tokens": 37,
               "prompt_tokens_details": {"cached_tokens": 0},
               "completion_tokens_details": {"reasoning_tokens": 40} },
    "cost": "0"
  }

cost 在**顶层**且是**字符串**。流式时还会单独发一帧:
{"choices":[],"cost":"0"}

## 此前丢了两样

1. completion_tokens_details.reasoning_tokens —— 输出里有多少是思考 token。
   没有它,客户端无法判断 completion_tokens 里多少是可见回答、多少是思考,
   而两者都按输出计费。
2. cost —— 唯一的费用信号,网关整个丢弃。Go 订阅是包月制恒为 "0",
   但 Zen 按量付费模型(以及未来的其它源)有信息量。

顺带修掉一处流式/非流式不一致:命中缓存时上游同时给
prompt_tokens_details.cached_tokens 和独立的 hit/miss,流式路径写成了 elseif,
只留 details,与非流式产出不同(只认独立字段的老客户端会看不到缓存)。

## 实现

- types.TokenUsage += CompletionTokensDetails;UnifiedResponse / UnifiedChunk += Cost
- opencodego/opencodezen 适配器映射两个字段;空 choices 帧改成 usage 与 cost
  都可带(早退只带 usage 会把同帧的 cost 丢干净 —— 新测试先抓到的就是这个)
- Gateway ChatCompletion / ChatChunk += cost,随终帧发(对齐上游的
  {"choices":[],"cost":"0"} 形态)
- Go 兜底 standardSSEChunk 同步支持(openai 系适配器不再漏 reasoning_tokens;
  纯 cost 帧不再被整体丢弃),新增 rawCostString 兼容字符串/数字两种形态

费用只做**搬运**:不解析、不换算、不汇总 —— 它是上游事实,且只有部分上游提供。

## 验证

经网关实测 gozen:deepseek-v4.1-flash,流式与非流式产出逐字段一致:
  prompt_tokens_details.cached_tokens=6784
  prompt_cache_hit_tokens=6784 / miss=148
  completion_tokens_details.reasoning_tokens=16
  cost="0"

测试:TestOpenCodeCostAndReasoningPassthrough(含「无数据不得凭空造字段」反例)、
TestOpenCodeStreamCacheFieldsMatchNonStream、TestTokenUsageMarshalsCompletionTokensDetails。

(cherry picked from commit c744ee151e)
2026-09-27 17:12:03 +08:00
d81074621a fix(opencode): 采纳客户端真实会话 id + 超窗消息不再被限流措辞封杀
两处都源于同一次排查:pi 到底有没有带会话标识、超窗为什么触发不了压缩。

## 1) 客户端会话 id:pi 一直在发,只是被配置关掉了

之前结论是「通用客户端不发会话 id」——只对了一半。pi 有会话 id,且能发:
pi-ai 的 createClient 在 compat.sendSessionAffinityHeaders 为真时,会把
平台会话 id(uuidv7,整个会话恒定)放到 x-session-affinity /
x-client-request-id / session_id 上。该开关默认 false,而 llmsproxy 的
provider 配置里没开,所以此前一直收不到。

现在网关按优先级采纳:x-session-affinity → x-session-id → session_id →
body 的 prompt_cache_key,并把值经 types.ChatRequest.ClientSession 传到
适配器 meta.client_session。适配器的会号种子优先级变为:
客户端会话 id > 首条 user 消息指纹 > 按源固定。

刻意不采纳 x-client-request-id:名字含 request,部分客户端每请求都换,
拿它当会话会让上游前缀缓存永不命中(pi 总会同时发 x-session-affinity,够用)。

实测:抓 127.0.0.1:8081 的真实 pi 请求,配置打开后收到
x-session-affinity = session_id = x-client-request-id = <子会话 uuid>。
上游缓存确为会话级隔离(同前缀、不同会号:A 冷→命中,B 首次仍为 0),
两个不同 header 值互不命中,反证网关确实采纳了客户端会话 id。

## 2) 超窗消息必须「干净」,否则被同链的限流措辞反向封杀

pi 的 isContextOverflow 先查 NON_OVERFLOW_PATTERNS(/rate limit/、
/too many requests/、Bedrock 前缀),命中就直接判为「非超窗」——**即使
消息里已经有 context_length_exceeded**,pi 也不会压缩重试。

而 AUTO 链的失败消息天生是多 tier 原因的拼接,超窗 tier(gozen 400
maximum context length)常与配额/限流 tier(429 token plan exhausted、
cooling、no free slot)同时出现。此前把 tier 明细原样拼在归一化标记后面,
等于让一条限流 tier 的措辞反过来封杀超窗识别。

现在超窗走独立的干净消息:
  context_length_exceeded: context window is full; reduce the length of
  the messages (gozen/deepseek-v4.1-flash)
只留超窗措辞 + 超窗源名,不带任何其它 tier 的文本。

测试:TestOverflowMessageSurvivesRateLimitedSiblingTier 用 pi 的完整判定
顺序(先 NON_OVERFLOW 后 OVERFLOW)断言同链限流 tier 不再封杀超窗识别;
TestClientSessionFromRequestHeaders / TestClientRequestIDIsNotUsedAsSession /
TestOpenCodePrefersClientSessionID 覆盖会话采纳与优先级。

(cherry picked from commit c9c09b2ba2)
2026-09-27 17:12:03 +08:00
6432b61eb6 fix(gateway): AUTO scope grants all models — restrict to routing mode only
A key with scope=[AUTO] could previously:
1. request ANY concrete model id directly (hasScopeModel/checkModelScope
   treated AUTO as a wildcard)
2. see the full 56-model list on /v1/models (intersectModels considered
   AUTO as grant-everything)

AUTO now only authorizes the AUTO routing mode. Direct requests to a
specific model require an explicit scope entry.

Also carries agentrouter.lua WAF fingerprint headers (Origin/Referer/
X-Requested-With) already staged on this branch.

Tests: TestHasScopeModelWithSourcePrefix updated; full suite green.
(cherry picked from commit 7fb8f96b82)
2026-09-27 17:12:03 +08:00
dd708c9680 fix(gateway): 上下文超窗错误归一化,客户端才能压缩重试
问题:上游返回上下文超窗时,客户端(pi)既不压缩也不重试,只看到一条
普通上游错误。链路两处叠加:

1. 措辞不在客户端识别列表里。pi 靠 @earendil-works/pi-ai 的
   OVERFLOW_PATTERNS(25 条正则)判断超窗,而 justworker 返回的是
   「请精简对话历史…(Context window is full…)」——与那 25 条一条都不匹配
   (最接近的 /context window exceeds limit/i 也不命中,因为措辞是 "is full"
   而非 "exceeds limit")。
2. AUTO 链路的每条 tier 错误按 80 字节 OneLine 截断,而 "Context window
   is full" 这类短诊断词常出现在尾部,正好被按字节切掉(直连路径是 160
   才保住)。

修法:
- 新增 overflowMarkers 识别超窗措辞(含中文写法),命中时把客户端可见
  错误归一化为 context_length_exceeded 前缀——它命中 pi 的
  /context[_ ]length[_ ]exceeded/i,超窗因此可被发现并触发压缩重试。
- AUTO 链路每条 tier 错误宽度 80→160,短诊断词不再被截断。

归一化只加前缀,原始诊断信息保留,便于定位是哪一层超窗。

测试:internal/gateway/overflow_err_test.go(5 例,含「标记必须命中 pi
正则」、非超窗不得误标、80 vs 160 宽度的回归对比)。
2026-09-10 20:38:39 +08:00
c309448414 feat(webui): Mono theme — pure white in light mode, pure black in dark mode
Adds a fourth accent alongside sakura/ocean/violet. Unlike those it is not just a
different hue: the coloured themes are glass surfaces (translucent cards with
backdrop-filter) floating over an animated gradient-mesh background, so setting
--card:#ffffff there still renders as a tinted grey. Mono therefore also switches
off the translucency and hides the blobs, so #ffffff is actually #ffffff and
#000000 is actually #000000, with greys carrying the hierarchy that hue carries
elsewhere. A side effect worth having: no backdrop-filter and no animated blobs
makes it the cheapest theme to render, which helps on weak GPUs and over remote
desktops.

Both light and dark variable blocks are defined, so the existing light/dark
toggle drives it with no extra wiring: light -> white, dark -> black.

Also fixes a latent theme bug found while checking contrast on black: the active
chart's grid baseline assigned the literal string "var(--line)" to
ctx.strokeStyle. Canvas 2D does not resolve CSS custom properties, so that was an
invalid colour the browser ignored, leaving the previous fillStyle (black) — an
invisible baseline on every dark theme. Colours used on a canvas now go through a
cssVar() helper.

Tests: TestUIThemeMatrix asserts every accent defines BOTH a light and a dark
block plus a picker button (a half-defined theme shows up as unreadable text, not
as an error); TestUIMonoThemeIsFlat pins the opaque surfaces and disabled blobs;
TestUICanvasColorsResolveVars fails if any ctx.strokeStyle/fillStyle is handed a
raw var().

Unrelated packaging fix in the same commit: dist:linux only built deb+AppImage
while build.linux.target listed rpm too, so `make gui-dist` silently skipped the
rpm that release builds are expected to produce. Makefile/README wording updated
to match.
2026-08-30 10:04:53 +08:00
5b6d6fc90d fix(gateway): keep the record ring small now that aggregates scan everything
The gateway still built Stats with NewStats(10000), a leftover from when the
ring WAS the source of the dashboard's numbers. With aggregates now computed
from the full audit history, a 10000-entry ring only means 10000 resident Req
structs — the live instance came up at 64 MB instead of ~40 MB, putting back
most of the memory the on-demand paging work removed.

The ring's only jobs are the status page's 5-minute SourceRecent /
SourceAverages windows and the dashboard's first screen, both of which fit in
defaultRingSize (500). Older records are paged from disk.

TestGatewayRingStaysSmall pins it so the constant cannot drift back up.
2026-08-30 09:34:25 +08:00
3ddae41f0c fix(gateway): aggregate the full audit history, repair record paging
Two problems reported after the on-demand log work landed.

1. Dashboard totals were wrong. LoadAudit only replayed the last 4 MB of the
   audit file, so requests/tokens/per-key rows reflected a window instead of all
   time — a regression in reported numbers, not just in presentation.

   The aggregates are now built by streaming EVERY audit file (oldest first, so
   the hourly quota buckets keep their intended trailing window) and keeping
   nothing per record: aggregate maps are keyed by key/model/source, so their
   size is bounded by cardinality. Measured on the production host: 29 MB /
   221k lines / 37k requests in ~260 ms at startup.

   What stays bounded is the RAW-record ring: a fixed-size reqRing keeps only the
   newest maxRecs records, so the ~25 MB that used to be spent appending every
   record into a slice is still saved. auditReplayBytes is gone, and
   replayPartial now means "an audit file could not be read", which is the only
   remaining way for the totals to be incomplete.

2. Scrolling to the bottom stopped loading more records. Two independent causes:

   * paintRecords rebuilt the entire table on every 5s poll whenever the row
     count did not exceed the first screen — the "is a paged view live?" test
     compared row counts and matched exactly on the first refresh — wiping loaded
     pages and resetting scroll position.
   * IntersectionObserver only fires on TRANSITIONS. With a short list, or after a
     page whose rows all duplicated the first screen, the sentinel stayed visible
     and never fired again.

   paintRecords now builds once (recsState.built) and later polls PREPEND only
   genuinely new rows; attachRecsObserver adds a scroll-position fallback;
   fillRecordsViewport loads until the list actually overflows; and
   loadMoreRecords chains (bounded) when a page yields no new rows, since the
   first fetch necessarily overlaps the first screen.

TestUIRecordsPagingWiring pins all four mechanisms structurally, since none of
them is reachable from Go. Test names/comments referring to bounded replay are
updated to describe the bounded RING instead, and both READMEs now state that
totals come from the full history while records are paged.
2026-08-30 09:29:42 +08:00
882288f67f fix(webui): send DELETE when removing keys and adapters
Deleting a gateway key from the admin UI did nothing and reported
"use GET /api/keys": delKey() called api() with an empty options object, so
fetch defaulted to GET and the request landed in the GET branch of
handleKeysAPI. delAdapter() had the identical bug and reported
"adapter code not exposed; edit in UI".

This is the third instance of the same mistake — ab20f1b fixed delSource and
delTemplate, missing these two — so it is now pinned by tests instead of by
review:

  * TestUIAPICallsDeclareMethod walks every api() call in the embedded
    index.html and fails if one passes an options object without a method
    (an AbortSignal-only read is allowed, being a deliberate GET).
  * TestUIDeleteHelpersUseDelete / TestUIMutatingHelpersUseWriteMethods pin the
    verb of each removal and write helper by name.
  * TestKeyDeleteRoundTrip covers create -> DELETE -> gone -> second DELETE is a
    clean 404, and TestCannotDeleteOwnKey keeps the lockout guard.

The 404 bodies for GET /api/keys/{key} and GET /api/adapters/{name} now name the
verb to use ("DELETE /api/keys/{key} to remove"), because that message is what a
mis-methoded client actually shows its user; "use GET /api/keys" read as though
the caller had done nothing wrong.

delSource's indentation, broken by ab20f1b, is also straightened out.
2026-08-30 09:07:05 +08:00
6575058556 feat(webui): scroll-paged record table, pool metrics, probe-aware health tags
Front-end half of the on-demand log loading plus the observability for the two
new scheduling mechanisms.

Records table:
  * the first screen comes from the dashboard poll, and an IntersectionObserver
    sentinel below the last row pulls the next page from /api/stats/records as
    the user scrolls.
  * rendered rows are capped at RECS_MAX_DOM=1000 (oldest rendered rows are
    dropped) so a long scroll cannot grow the DOM without bound, with a "back
    to newest" button to reset cheaply.
  * releaseRecords() drops the buffer, disconnects the observer and aborts the
    in-flight fetch (AbortController) on tab switch, on key-filter change and
    on pagehide/beforeunload — leaving the page releases everything at once.
  * the poll no longer rebuilds the table once extra pages are loaded, so the
    5s refresh cannot throw away scrolled history.
  * when the server reports replay_partial, the filter row states that the
    aggregates cover the recent audit tail and points at CSV for full history.

Adapters page shows each pool as created/max plus in_use/idle and the live
grow/shrink steps, with the sizing rule in the hover text.

Priority page health tags now distinguish cooling / probe-ready / probing
instead of a flat "cooling", and the tooltip spells out when the window opened,
when the single probe is allowed through and when the slot clears completely.
core.AutoSlotState carries cooldown_from / probe_after / probing / probeable to
feed this.
2026-08-30 08:06:11 +08:00
d42c02b15d perf(gateway): load request logs on demand instead of holding them in memory
Startup RSS on this deployment was 56 MB with a 29 MB audit log and ~10 MB
without one: LoadAudit() json-unmarshalled the ENTIRE file into the aggregates
and kept a 10000-entry ring of raw records. Two more paths had the same shape —
AuditRecords() materialized a whole export window into a []Req before sorting
it, and a dashboard poll serialized the full ring so the browser could render
300 rows of it.

The audit file is now the source of truth and memory only holds the live
window:

  * LoadAudit replays only the last auditReplayBytes (4 MB) and drops the
    truncated first line; the ring default drops 10000 -> 500, which still
    covers both of its consumers (the status page's 5-minute SourceRecent /
    SourceAverages windows and the first screen of the records table).
    replayPartial is exported so the UI can say the totals cover a window
    rather than all time. Token-quota accounting is unaffected: it reads the
    modelHour buckets, not the ring (pinned by a test).
  * AuditPage(cursor, limit, key) pages records straight off disk, reading the
    newest file backwards in 64 KB chunks and returning as soon as the page is
    full. The cursor is "<file>:<offset>" and walks into rotated .old files;
    a cursor whose file rotated away reports rotated=true so the client can
    reset instead of silently skipping records. No state is cached between
    requests and the file handle is closed before responding, so "release when
    the user leaves the page" is guaranteed by never retaining anything.
  * StreamAuditRecords(from,to,key,fn) replaces the accumulate-then-sort export
    path; the CSV handler writes rows as they are read and flushes every 1000,
    and a write error (client gone) aborts the walk. Export memory is O(1)
    regardless of the window. AuditRecords is kept as a test-only wrapper.
  * Snapshot ships one screen (firstScreenRecords=100) by default; aggregates
    are untouched.
  * Audit rotation 64 MB x 10 -> 16 MB x 16: same 256 MB total budget, but a
    smaller newest file keeps the first reverse page cheap.

New route: GET /api/stats/records?before=&limit=&key= (non-admins are pinned to
their own key by exportKey). /api/status additionally reports adapter_pools for
admins.

Measured with production's 29 MB audit copied to the test instance: startup RSS
19.0 MB (was 56 MB); scrolling 10 pages (1000 records) +0.7 MB; exporting the
full history (36441 rows / 4.4 MB CSV) +0.1 MB with no residual growth.
2026-08-30 08:05:54 +08:00
2ed1f0ecde style: gofmt the tree
gofmt -l reported 13 files with misaligned struct tags / stale formatting.
This commit contains ONLY formatting: no behaviour change, no logic touched.
Files that also carry real changes in this series are formatted by their own
commits.
2026-08-30 08:04:22 +08:00
624fd74b45 fix: anthropic tool-call round-trip, cache zero-hit parity, round-robin load balancing
anthropic.lua v3.0.0:
- Issue 1: tool_result/tool_use round-trip
- Issue 3: thinking default OFF (opt-in via extra_body.thinking)
- Issue 4: tool_choice mapping
- Issue 5: collect_blocks preserves unknown part types
- message_stop no longer emits done=true (was overwriting tool_calls finish_reason)
- cache_read_input_tokens normalized even at 0

gemini.lua:
- transform_response was missing cachedContentTokenCount

openai.lua (Issue 6):
- transform_error handles flat envelopes, nginx HTML, bare text

chat.go mergeUsage:
- Keep PromptTokensDetails even when CachedTokens=0

scheduler.go:
- Remove sort.SliceStable by Pref; round-robin cursor is the only LB mechanism

provider.go ModelAvailable:
- Also check Pref() > prefMin, persistently failing slots exit cands

presets.go:
- 17 built-in source templates

Tests: 6 new test functions, 2 updated for new semantics
2026-08-28 12:02:46 +08:00
ab20f1be48 fix(webui): specify DELETE method for source/template removal
delSource and delTemplate called api() without a method, so fetch
defaulted to GET; the DELETE handlers never ran and the UI silently
left the item in place.
2026-08-27 12:14:41 +08:00
8334cffbc9 feat(templates): source templates for multi-key balancing
A template stores every source field except name and api_key, so
operators spin up N key-bearing sources from one shared skeleton
instead of duplicating the whole source block N times.

- config: SourceTemplate type + RuntimeConfig.SourceTemplates stored
  in runtime.json alongside runtime sources
- store: UpsertTemplate / ListTemplates / RemoveTemplate
- core: Templates / SaveTemplate / RemoveTemplate
- gateway: GET/POST/DELETE /api/source_templates
- webui: source list gains a Templates button opening a manager with
  per-template edit/delete; the add-source dialog gains 'from
  template' (event-delegated picker) and 'as template' (card modal)
  buttons in its header; z-index fixed so the template editor layers
  above the manager
2026-08-27 12:09:38 +08:00
a42ff62d06 feat(scheduler): separate image-generation AUTO chain with UI toggle
Image models previously could not be scheduled through a priority
chain: the chat AUTO chain explicitly skips image-kind slots, and
AUTO image requests fell back to unordered registry discovery.

- config: add auto_image rules (auto_image yaml / image_rules json);
  legacy auto rules keep their meaning as the chat chain
- core: buildAutoImageChain mirrors buildAutoChain with inverted kind
  filter (image-only); SaveAutoImageRules + AutoImageRules/AutoImageChain
- scheduler: ChainImage walks the chain tier-by-tier with round-robin
  and preference ordering, skipping cooling slots
- gateway: handleImage AUTO now runs down AutoImageChain when one is
  configured (falls back to legacy discovery otherwise) and records
  the actual served model; handleAutoAPI GET returns image_rules and
  PUT accepts image_rules independently of rules
- webui: priority page gains a chat/image toggle editing two
  independent lane sets; add-slot picker filters by active kind;
  persistAuto writes only the active chain's field
2026-08-26 21:03:51 +08:00
66585549f1 fix(webui): allow image models in priority/AUTO chain editor
The priority page and key-scope pickers skipped kind=image models,
so image sources (e.g. Kwai-Kolors/Kolors) could not be placed in
the AUTO chain or granted per-key. The backend already supports it:
handleImage resolves AUTO through the chain then filters with
imageOnly, while chat requests are protected by chatOnly, so image
slots never receive chat traffic.

Image models now appear in the priority canvas, the add-slot picker
(labelled ' (image)'), and the key scope dialog.
2026-08-26 20:36:05 +08:00
747dff5b76 fix(gateway): record actual served model for image requests
handleImage recorded the raw request model id, so AUTO image
generations showed up as model=AUTO in the request records and
by-model aggregates instead of the image model actually served
(e.g. Kwai-Kolors/Kolors).

UnifiedResponse gains an optional Model field; Provider.Image fills
it with the resolved id (AUTO resolves to the source's best image
model), and handleImage prefers it when writing the audit record.
2026-08-26 20:28:20 +08:00
e48baa1bb3 fix(webui): missing HTTP method on fetch calls with body
6 call sites used the fetch default GET while sending a request body,
which browsers reject outright ('Request with GET/HEAD method cannot
have body'). Affected flows: create key, save source, save auto rules,
upload adapter, update key model scope, and /api/chat streaming.

All now send the method their backend handlers require (POST or PUT)
with an explicit Content-Type.
2026-08-26 00:34:57 +08:00
dev
77b00bfe3f feat(gateway): brute-force protection for /api/login
Prerequisite for removing the nginx global-auth layer in front of the
gateway: llmsproxy must defend its own login endpoint.

Design:
- per-IP failure counter: 5 consecutive failures trigger an exponential
  lockout (30s base, doubling per extra burst, capped at 30min); 15min of
  quiet forgives the counter
- global budget: max 100 failures/minute across all IPs so a distributed
  spray cannot outrun per-IP windows
- locked-out and over-budget attempts get the SAME 'invalid gateway api
  key' 401 as normal failures — no oracle to probe lockout state, no info
  leak on key validity timing
- successful login clears the IP's counter entirely
- clientIP(): prefers X-Real-IP (trusted nginx proxy), falls back to
  RemoteAddr host

Tests: lockout engages at threshold with identical replies, valid keys
rejected while locked, other IPs unaffected, success resets counters,
X-Real-IP extraction.
2026-08-25 11:01:49 +08:00
dev
21ec8f59d8 feat: surface zero cache hits — distinguish 'missed' from 'not reported'
Live testing across the zen pool showed models report
prompt_tokens_details.cached_tokens even when the hit count is 0 (e.g.
nemotron-3-ultra-free returns cached_tokens:0, audio_tokens:0,
cache_write_tokens:0). The previous >0 guard dropped those objects, so a
cache-enabled upstream looked identical to one without cache support.

- types: PromptTokensDetails.CachedTokens always emitted (drop inner
  omitempty) so clients see cached_tokens:0 explicitly; dsh reads it as
  a 0% hit instead of 'no data'
- adapters (9): forward prompt_tokens_details whenever the upstream
  provides it (presence check instead of >0)
- Req: add cache_reported flag set when usage carried cache accounting;
  WebUI shows an amber 0% tag for reported-but-missed rows and keeps
  the em-dash only for sources that never report cache data
2026-08-25 10:04:44 +08:00
dev
24609289e8 feat: record cache hit/miss per request in audit trail and WebUI
- Req: add CacheHit and CacheMiss fields (carrying upstream cache
  accounting from either prompt_tokens_details.cached_tokens or legacy
  prompt_cache_hit_tokens)
- recordChatUsage (non-streaming): copy cache fields from resp.TokenUsage
- pumpStream (streaming): write lastUsage cache fields back onto rec at
  stream end, so streaming requests carry cache data too
- CSV export: add first_byte_ms, cache_hit_tokens, cache_miss_tokens
  columns alongside the existing latency/prompt/completion
- WebUI request-records table: add a Cache column showing hit% per row
  (green/amber tag with tooltip hit/miss breakdown; em-dash when the
  upstream reported no cache data)
2026-08-25 09:36:21 +08:00
dev
18c2385c61 feat: add per-source TTFB and tokens/s metrics to status page
- Req: add FirstByteMs field (ms to first byte, tracked for streaming)
- Stat: add FirstByteSum for aggregation
- SourceAverages(): new method computing per-source avg TTFB and tokens/s
  from the in-memory ring (300s window)
- SourceStatus: add AvgFirstByteMs and AvgTokPerS fields
- pumpStream: record FirstByteMs after first SSE chunk sent to client
- singleChat/singleChatAuto: set FirstByteMs = LatMs (non-streaming)
- handleStatusAPI: populate the new SourceStatus fields from SourceAverages()
- WebUI source table: two new columns showing TTFB (s) and Tokens/s
2026-08-25 08:24:08 +08:00
dev
045ecf47bc feat(types): pass through upstream cache tokens in TokenUsage
dsh displays cache-hit %, but llmsproxy dropped every upstream's cache
fields — deepseek prompt_cache_hit_tokens, OpenAI prompt_tokens_details.
cached_tokens, anthropic cache_read_input_tokens, gemini cachedContentTokenCount.

Changes:
- TokenUsage: add PromptTokensDetails (with CachedTokens) + PromptCacheHit/Miss
- MarshalJSON: emit prompt_tokens_details.cached_tokens (OpenAI v2 standard)
  and prompt_cache_hit/miss_tokens (DeepSeek legacy) — dsh reads the former
  first, falls back to the latter
- mergeUsage: preserve cache fields across stream chunks
- standardSSEChunk: parse the upstream raw prompt_tokens_details too
- deepseek.lua: forward prompt_cache_hit/miss_tokens + create
  prompt_tokens_details from them
- openai.lua: forward prompt_tokens_details.cached_tokens and legacy
  prompt_cache_hit/miss_tokens; normalize legacy hits into the standard
  object so dsh sees them regardless of upstream format
- anthropic.lua: map cache_read_input_tokens → prompt_tokens_details
- gemini.lua: map cachedContentTokenCount → prompt_tokens_details
2026-08-25 07:57:32 +08:00
dev
334b984c25 fix(webui,gui): repair dead export button + wire chat clear; prune UI redundancy
WebUI (internal/gateway/ui):
- BUG: the export modal's custom-range button called
  downloadStatsCsvFromForm() which was never defined — clicking it threw a
  ReferenceError and nothing downloaded. Implement it: reads #exp-from /
  #exp-to date inputs and forwards to downloadStatsCsv.
- BUG-adjacent: clearChat() existed but was reachable from no control —
  add a Clear button to the chat composer so conversation reset is actually
  possible (+ cClear i18n zh/en).
- remove byte-identical duplicate html[data-theme=dark] CSS block (15 lines)
- remove 8 dead CSS rules (.keys-grid .m-model-row .scope-add/.scope-box/
  .scope-chips .scr-blocks .tag-warn .twrap) and the never-consumed
  --accent custom property
- remove 3 dead JS functions (activeTab/findSlots/scopeUncomb; lastTab decl kept)
- remove 24 dead i18n keys x zh/en (~55 lines) — legacy of the replaced
  key-scope editor, matching the removed .scope-* styles

GUI (cmd/gui/main.js):
- BUG: stopCore() set app.isQuitting=true and nothing reset it — after using
  tray 'stop core', closing the window quit the whole app instead of hiding
  to tray, and core crash auto-restart stayed disabled. isQuitting now only
  flips in restartCore (scoped) and before-quit.

Verified: go vet/test green; node --check on all three GUI js files and the
WebUI inline script.
2026-08-24 23:13:54 +08:00
dev
f3d5ba6cea refactor(gateway): deduplicate stats CSV export paths
- extract csvHeaders() (Content-Type + Content-Disposition) shared by both
  export branches
- extract exportKey(): the identical admin/user key-filter logic existed
  twice (JSON path + keys-csv); now all three call sites share one function
- inline the nine single-use intermediate variables in the keys-csv loop
2026-08-24 22:41:48 +08:00
dev
98847adfd6 refactor(gateway): collapse chat.go four-way duplication
singleChat / singleChatAuto / streamChat / streamChatAuto shared ~260 near-
identical lines (diff after stripping comments was empty). Extract three
shared bodies and shrink all four entry points to thin dispatchers:

- failChat: error → record + writeError, shared ChainErr extraction
  (errors.As returns false for direct-path errors, so the two paths stay
  equivalent without a branch)
- recordChatUsage: exact upstream numbers win, byte estimates fill gaps
- writeChatCompletion: unified ChatCompletion rendering (model name is the
  only direct/auto difference, passed in)
- pumpStream: the full SSE pump (preamble, delta loop, terminal finish +
  usage chunk, [DONE]) — shared by both streaming entries

Behavior change (pinned by TestDirectStreamFailoverAuditSource): direct
streams now pin rec.Source to the source that actually served the stream
after a failover, instead of discarding it (_, usedModel). The audit row
previously recorded the first candidate, which was wrong on failover.

chat.go: 982 → 896 lines (−86).
2026-08-24 22:38:38 +08:00