Commit Graph

48 Commits

Author SHA1 Message Date
d0ddc9754d feat(billing): 计费页支持按日/周/月/全部,与统计页同口径
宿主统计页已支持周期,计费页还是终身累计 —— 同一个问题在两个页面重复出现,
且两个页面口径不一致本身就是错。

## 为什么要补日级维度明细

原来只有 total 和 by_day 带时间维度,by_source/by_model/by_key 是终身累计。
只改 total 的话,页面会显示"今日开销 $0.05",下面三张表还是全量数据 ——
数字对不上,而这正是周期视图要消除的错配。所以在写入时按天折叠成本
(by_day_src/model/key)。成本在写入时就已定价,浏览器只做求和,不重新
定价,显示金额不会与持久化金额漂移。

按天为键的表不随流量增长(一年 365 项/维度),所以不设裁剪。

## 口径

UTC,与网关 dayKey 和 /api/stats 的周期窗口一致 —— 峰谷小时不会在费用
视图和用量视图落到不同一天。周为 ISO 周(周一起)。降级/未定价没有日级
计数,周期视图下**隐藏**这两张卡而不是显示终身值,那正是要消除的错配。

## 判据(7 条 + 6 个变异)

判据断言**渲染后的 DOM**,不走 IIFE 内部函数:早先版本加了测试钩子去直接
调 daysInWindow/rescale,结果 harness 里的假 Date 先后两次出问题(整体替换
构造器破坏 toISOString;子类化导致 getUTCDay 返回 NaN),症状都是
"Invalid time value",看起来完全像产品 bug。

变异验证抓到判据本身的缺陷,值得一提:
- 捕获「总账不累加」和「周起点改周日」两个变异时**全部判据仍绿**——因为
  我断言的是 by_source 表格,而它由维度表独立算出,跟 total 无关。
- 补了 KPI 断言后又漏放「维度折叠不求和」——"页面里存在 60" 太弱,
  60 同时出现在 KPI 和 token 列里。改成锚定 srcA 自己的金额单元格
  (USD 60.0000),last-day-only 会是 30,才抓得住。

判据报错时先怀疑判据——这次三次都是判据的问题。
2026-10-02 12:41:35 +08:00
d1da40e493 fix(billing): 页面文字溢出卡片 + 数字千分位
用户报 billing 页"文字超出展示框"。真因是三处叠加,都不是文案问题:

1. KPI 卡是 CSS grid,grid item 默认 min-width:auto。2e8 级的 prompt
   tokens(实测 270,149,209)无法收缩,于是把 grid 轨道撑出容器 → 整页横向
   溢出。修法是 min-width:0 + overflow:hidden(不给 min-width:0 的话,
   grid 子项永远不肯收缩,这是 grid 最常见的溢出坑)。
2. 表格写死 min-width:560px,窄视口下必然溢出。改 width:100% +
   table-layout:fixed,列宽由布局分配而不是由内容撑开。
3. 长名称(deepseek-v4.1-flash 这类模型 id)撑宽单元格。名称列改 ellipsis +
   title 悬停看全名;数字列 word-break:break-all 在列宽内换行。

顺手:所有 token/请求数走 toLocaleString 千分位。原始 9 位数字读起来要数
零位数,分组后 270,149,209 一眼可读,也顺带缩短了字符串宽度。

CDP 实测(820px 窄视口,逼出溢出条件):
  main/body 横向溢出 = no
  table=338 < card=366(修复前 min-width:560 必然 > 366)
  KPI 输入 tokens = 270,149,209
  billing 页无控制台错误

残留(不影响布局):表头 TH 在固定布局下 48>42 轻微超出自身格,因为
word-break 不拆单个长词;表格整体仍在容器内,未产生页面滚动。

全量测试全绿(8 包)。
2026-10-02 12:16:40 +08:00
aa10ee8c27 fix(billing): 插件页不可达(真·空白根因)+ i18n + 溢出 + 状态页 tile
## ★ 用户报告「Billing 页还是空白」—— 上一轮的验证有漏洞
上一轮我用 goTab('billing') 直接调用验证,显示"有数据、无错误"就下了结论。
但用户是**点侧栏按钮**。真实点击路径走 goTab,而 goTab 只遍历硬编码的 TABS
常量来切换 `hidden` 类 —— 插件页不在 TABS 里,所以 #tab-billing 的 hidden
**永远不会被移除**。内容一直躺在 DOM 里(KPI/表格都填好了),只是不可见。

这个 bug 没有任何报错:注入正确、数据正确、API 200,唯一的问题是宿主页的
路由逻辑没把插件页纳入。而它是上一轮「TABS 收敛为单一常量」时留下的:
收敛让三处共用一个常量,却没让插件页进入它。

修法:goTab 同时遍历 PLUGIN_PAGES。PLUGIN_PAGES 从 const 改为 var 并**提前到
goTab 之前声明** —— const 在文件后部声明的话,goTab 的读取落在 TDZ 里,
第一次点击插件页就会抛 ReferenceError(同类问题这个文件里已是第二次)。

判据 TestPluginPagesAreReachableByGoTab 锁两件事:goTab 遍历插件页集合 +
声明在使用之前。变异验证:删掉遍历 → 红;var 改回 const → 红。

## 插件 UI 不跟随多语言
宿主的 applyI18n/data-i 只覆盖**宿主渲染的标记**;插件注入的 HTML 对它不可见,
所以整个 UI 切中文时 Billing 页还是英文。

pluginAPI 增加 lang(getter,实时值)与 onLangChange(切换回调)。
billing 页所有文案改走双语字典:KPI、表头(fresh/cache/cache%)、区块标题
(占位后由脚本填)、空态、状态页 tile 标签。切换时立即重渲染标题,
不用等下一次 fetch。

## 部分页面超出 UI 区域
#main 只有 overflow-y,插件页内容(8 列表格 min-width、长字符串)会横向撑破。
两层修:插件 pane 统一 min-width:0/max-width:100%/overflow-x:auto(第三方
任意 HTML 的兜底,与原生 pane 一致);billing 的宽表格在自身容器内滚动。

## 状态页 tile 的 TypeError(每次重绘都报)
tile 的 tick() 在 await 之后直接 getElementById(...).textContent = ...,
但状态页每次刷新都整体重建 pane,元素可能已不存在 → null 属性赋值。
await 之后重新取元素并判空。

## 顺手补的缺口
上一轮加了缓存表格列,但 KPI 卡片漏了(那次替换 assert 失败后重试只重做了
表格)—— 缓存命中率在表格里有、KPI 里没有。本次补上。

## 验证
真实浏览器(禁缓存、真实点击侧栏按钮):pane 可见、KPI 9 项、表头双语、
语言双向切换正确(Per source ⇄ 按源)、无水平溢出、无 billing 控制台错误。
生产数据:Total USD 0.566798 / 732 请求 / 降级 231 / 2.09 亿 prompt tokens。
387+ 测试全绿。

## DSL(进行中,未完)
config.BillingDSL(active + profiles + rules,rule 按 url 匹配 mode=free/
token/subscription/unpriced)与 internal/billing.Compile(url 规则 → 插件
prices 表,含峰谷窗口的形状编译 —— 之前手写 JSON 两次弄错的正是这个形状)
已落地并通过校验/编译;core 启动接线已写。profile 切换 API 与 WebUI 选择器
未做,生产 config.yaml 也尚未写 billing 段 —— 下一轮继续。
2026-10-02 11:47:24 +08:00
fbdf0dea10 fix(billing): 缓存命中统计缺失 + Billing 页空白 + 侧栏图标
三个问题都来自生产实测,不是代码审阅。

## 1. 缓存命中被计费却不被统计
网关确实从上游 usage 提取了 prompt_cache_hit_tokens(审计里能看到
cache_hit_tokens: 270104 / cache_reported: true,占 prompt 的 99.9%),
costFor() 也用它给缓存段定价了 —— 但**没有任何 bucket 记录它**。
结果:一个 99.88% 命中率的网关,报表显示 prompt_tokens 却看不出其中
多少是缓存读,也无从按源/模型/key 看命中率。

每个 bucket 现在多三个字段:
  cache_hit_tokens    命中数(按上游上报)
  cache_fresh_tokens  未命中的 prompt
  cache_reported_reqs 上游确实上报了缓存数的请求数

第三个字段是刻意的:**「零命中」与「上游根本不上报」在命中总量里完全一样**,
而它们在「缓存折扣有没有生效」这个问题上含义相反。没有它就无法区分,
只能猜。

chat.go 的 payload 之前**没有** cache_reported(审计有、插件没有),
所以任何插件侧的缓存统计都只能猜 —— 已补上。

旧 state 文件的 bucket 没有这些字段:Lua 里 nil + number 会抛错,而钩子抛错
会让**该请求完全不记账**(一个统计缺口会变成静默缺口)。add() 里做了回填。

UI 增加 fresh/cache/cache% 三列 + Cache hit rate KPI;未上报的显示 n/r 而不是 0%。

## 2. Billing 页空白:render() 引用了未定义的 s
`render(st)` 里两处 KPI 写成 `s.degraded_reqs`,ReferenceError 让整个渲染
中断,所有表格停在初始的空 innerHTML。症状是「页面加载了但什么都没有」,
而 /api/plugins/billing/state 返回 200 且有真实数据 —— 载荷完全正确,
DOM 是空的。

更糟的是 refresh() 里的 `catch (e) { /* never break the page */ }` 把错误
**静默吞掉**了:网络面板一切正常,页面什么都没有。现在 catch 会
console.error(仍然不抛,装饰性组件不该拖垮宿主页,但必须留痕)。

## 3. 侧栏图标
billing 声明 icon = "💰",而原生 tab 全是内联 SVG(stroke: currentColor)。
emoji 尺寸不对、不跟随主题。

WebUI 增加 pluginIconHTML:插件图标可以是文本,也可以是内联 SVG。
**SVG 走严格白名单**(tag + 属性都是 allowlist,不是 denylist)——
插件是在运维者浏览器里跑的第三方代码,不能"信任插件";但也不能直接拒绝
SVG,因为那是唯一能和原生 tab 视觉一致的方式。

用真实 Chromium 验证 12 个用例,全部挡住,包括 foreignObject 里嵌 HTML
命名空间 <img onerror> 这个经典绕过(整体丢弃,所以 img/onerror 也没了)。
★ node 里没有 DOMParser/jsdom,所以没法在单测里跑这个过滤器 —— 用正则近似
会得到一个"测试通过但浏览器里失效"的过滤器,这比没有测试更糟。

顺带修了过滤器的两个真缺陷:输出里嵌套了空 `<svg></svg>`,且 viewBox
是从包装元素读的(永远是 null)而不是插件自己的,所以任何自定义 viewBox
的图标都会丢失。

## 判据(新增 7 项,全部变异验证)
写「注入脚本能否正常执行」这个守卫时我错了四次:
  1. 静态扫「已声明的名字」→ 把 HTML 字符串里的 CSS 类名(class/div/td)
     全报成未定义
  2. 用 CSS 选择器解析器查样式表 → 报样式表本身坏了
  3. 只挂 process 的 uncaughtException → 脚本在 IIFE 里异步跑,错误是
     unhandledRejection,判据对原 bug 全绿
  4. 只查「有没有抛错」→ render() 开头是 `if (!st) return`,传错字段是
     **静默 no-op**:不抛、不打日志、不报错,只是页面空白
最终判据是:在 node 里用 DOM stub 真跑一遍,同时要求「无异常」且
「至少写进一个容器」,并监听 console.error。变异验证:还原 s → 红;
render 收到 undefined 字段 → 红。

表头/行列数一致性也有守卫:row() 加了缓存列而表头没加时,表格会整体错位
(cache% 落到 completion 列下)—— 渲染正常、有数据、但要仔细看才发现。

## 生产验证
重启后价目表与累计账完整保留(1.17 亿 prompt tokens)。
新请求缓存统计生效:cache_hit 947,436 / cache_fresh 888,
cache_reported_reqs 7 / 395(其余来自旧 state,正是该字段存在的意义)。
真实浏览器:表格 3 行、KPI 7 项、表头 name/cost/reqs/prompt/fresh/cache/cache%/completion、
SVG 图标 currentColor 渲染、控制台无 billing 错误。391 个测试全绿。

## 另发现一个无关 bug(未修)
首页 stats 图表抛 IndexSizeError: arc 半径为负(-2),在 ui/index.html 的
paintStats 附近。属状态页图表,不在本次范围。
2026-10-02 11:16:15 +08:00
30064696b3 perf(plugin): 去掉钩子热路径的 JSON 往返 + 同 stage 跨插件并行
## 1. 去掉 JSON 往返(快路径)
实测单次 Fire 14.6µs,其中 json.Marshal 4.0 + json.Unmarshal 5.6 = 9.6µs,
**67% 花在把 map[string]interface{} 序列化再反序列化**,而紧接着的
pushGoValue 本来就能直接遍历这两种类型。改为按类型直接转换(fastvalue.go),
只对不认识���类型才回落 JSON —— 陌生字段仍然会被送到插件,而不是消失。

快路径与 JSON 路径逐字节等价由 TestFastPathMatchesJSONPath 锁住(8 组载荷,
覆盖 int/uint/float 各宽度、嵌套、slice、map[string]string、未知类型)。
还有一条专门防止「优化悄悄失效」:TestFastPathIsActuallyUsed 用真实的
request_end 载荷断言它确实走快路径。

同一份代码 A/B 实测:JSON 往返 56.4µs → 快路径 35.3µs(省 37%)。

## 2. 同 stage 跨插件并行
参照 /home/program/TrueAgent 的 StageHost.RunStage:
  - **快照后释放锁**再并行 —— 它记录过一次自死锁(p.Stop → onExit → ReclaimOwner
    要拿 registry 锁,持锁并行即死锁)。这里同理:钩子可能经 admin API 增删插件,
    那条路径要拿 ps.mu 写锁,所以并行段内不持任何 ps 锁。
  - 每个 goroutine recover。
  - 单插件走直连路径,不付 goroutine 代价(生产就是这种配置)。

**与 TrueAgent 不同的一点**:它可以放心并行,因为 handler 只写 ctx.Response 并有
IsResponded() 仲裁;我们的钩子返回 table 会合并进 payload,而
docs/plugins.md 明确承诺「payload 原样传给下一个插件」。所以合并**按插件加载
顺序**执行,结果确定,不依赖调度;代价是钩子之间不再互相可见 —— 这是一处
**契约变化**,已在文档里写明,并说明随核心发布的 billing 从不返回任何值
(代码注释就写着 "nobody downstream would read a return value")。

实测收益(真实二进制,三实例对照,3000 请求):

              无插件     1 插件      4 插件
  稳态并发32    849 rps   768 (-9.5%) 741 (-12.7%)
  突发并发64   1524-1893  1182-1676  1064-1443

4 插件只降 10-20%,而并行前实测 4 插件是 63.8µs vs 单插件 14.6µs(-300%)。

## ★ 我自己造成的两次性能事故
**① 持久化把热路径拖慢 26 倍。** 最初的快照在钩子路径上做:走 luaValueToGo +
json.Marshal + json.Unmarshal 三重转换,每请求 264µs,Fire 从 14.6µs 变成 385µs。
改成 saver 按自己节奏拉取(钩子只标记 dirty,flush 时才快照),385µs → 25.7µs。
**这里还踩了第二次 use-after-free**:让后台 goroutine 去读 Lua 表,vm.Stop() 后
那是已释放内存(SIGSEGV)。安全性现在由「Plugins.Close 等 saver 的最后一次
flush 完成后,调用方才停 VM」保证。

**② 基准被自己的后台写入污染。** 关掉 markDirty 反而测出 36µs、比开着还慢,
方向完全反了 —— 是 saver 每 2 秒写盘混进了计时。加了 DisableStatePersistence
后数据才可信。

## 判据(11 项,全部变异验证)
快路径等价/确实生效/不别名输入 + 并行与单插件路径合并一致 + 合并顺序确定 +
抛异常的钩子不拖累同伴 + 每插件恰好执行一次 + 并发 Fire 安全 + Fire 期间不持
注册表锁 + 真实 billing 在并行下正常 + 持久化 7 项。

变异:改坏合并顺序 → 红;去掉单插件路径的合并 → 红(3 个既有测试同时抓到)。
★ 「删掉 recover」这个变异**没有**让判据变红,查下去发现 golua 把 error()、
nil 索引、调用 nil、深递归全部转成 error RETURN,不产生 Go panic —— 那个测试
根本没测到 recover。已改名 TestThrowingHook 并在注释里写明 recover() 当前无法
被 Lua 触达,保留它是为了守 Go 侧。留一个「看起来有覆盖」的断言比没有更糟。

## 端到端(真实二进制 + 真实 billing)
20 万请求全 200,rps 1870,p99 96ms,RSS 37.9→42MB 有界;
负载停止后四个插件计数**完全一致**(231745),hook_errors 为空;
systemctl restart 后 billing 仍是 231745 —— 并行与持久化同时生效。
381 个测试全绿,含 -race。
2026-10-02 10:45:20 +08:00
ce2032435a fix(plugin): 插件 state 持久化 —— 重启不再丢账
## 问题(压测实测)
插件 state 活在 Lua VM 里,进程一死就没了。实测线上量级:
  重启前 {"requests":218241,"prompt_tokens":26188920,...}
  重启后 {"requests":0,"prompt_tokens":0,...}
对计费插件来说这不是舍入误差,是功能本身没生效 —— 它存在的意义就是那个
不断累加的数字,而一次 systemctl restart 就能把它抹掉。

## 设计
- prices(配置)与 state(累计历史)**分开存**在同一个文件里但不同字段。
  SetState 在内存里已经这么分,磁盘必须同意:合并会让改价看起来像清零,
  或让恢复历史时顺带复活过期价格。
- 原子写(临时文件 + rename):写一半崩掉时上一份仍可读,而不是留下一个
  解析失败的半截 JSON —— 那等于这次也丢。
- 损坏文件只警告不阻断启动。转发不能依赖插件的账本活着。
- 防抖后台刷:钩子路径只标记,真正的写在一个 goroutine 里合并进行。
  计费插件每请求都改 state,同步写会把一次 JSON 编码 + 文件写放到热路径上
  (实测钩子本身已经 14.6µs,写会盖过它)。
- Core.Close 必须先刷插件再停 VM:flush 要读 Lua 表,vm.Stop() 之后读的是
  已释放的内存。

## ★ 实现中踩的四个坑(都由测试或崩溃直接暴露,不是推测)
1. **后台 goroutine 碰 Lua = use-after-free**。最初让 flush 线程去读 Lua 状态,
   vm.Stop() 后那是已释放内存 —— 表现为 golua 里的 SIGSEGV,不是干净报错。
   改成:钩子路径(VM 必然存活、已持 p.mu)取快照,后台只写文件。
2. **自死锁**:markDirtyLocked 被 invoke 调用,而 invoke 全程持 p.mu,
   再 Lock 一次就是死锁。lua 包测试直接挂到超时。
3. **luaToJSON 独占整个栈**(每条路径结尾都 SetTop(0))。连续调两次读两个
   字段时第二次访问的是不存在的槽位 —— 这个绑定不 panic,直接 SIGABRT。
   改为每次重建栈。中间还因为提前 return 没 Pop 而让栈逐次错位。
4. **快照顺序**:先快照后读返回值,会把钩子的返回值清掉,于是每个"有意见"的
   插件静默变成"没意见",而文档承诺的"返回 table 合并进 payload"就废了,
   且没有任何报错。

## 判据(7 项,全部变异验证过)
重启后总计保留 / prices 与 state 分离 / 纯 prices 更新也持久化 /
钩子返回值不被快照吃掉 / 损坏文件降级不阻断 / 500 次变更合并成个位数次写 /
Close 刷出尾部。

变异结果:
  关掉 mark          → TestStateSurvivesRestart + TestPricesAndStateAreSeparate 红
  Close 不等 flush   → TestCloseFlushesTail 红
  prices-only 不写盘 → TestPricesOnlyUpdatePersists 红
  还原快照顺序       → TestHookReturnValueSurvivesSnapshot 红
★ 第一次跑「关掉 mark」时判据没报错,原因是我的变异脚本写出未使用变量导致
  编译失败 —— go test 根本没跑测试,我却读成了"通过"。换成 _, _ = 后如期变红。

## 端到端
隔离实例发 12 次请求 → systemctl restart → requests 仍为 12,token 数不变。
371 个测试全绿。
2026-10-02 10:17:07 +08:00
a78f7cb6c5 feat(plugin): 启用/禁用 + 磁盘列表 + 峰谷定价 + 随核心发布
## 插件管理后端
- PUT /api/plugins/{name} {"enabled":bool}   启用/禁用
- GET /api/plugins/{name}                    读源码(编辑器用,与 /state 区分)
- GET /api/plugins 的 on_disk 字段            列出目录里所有 .lua 及其加载态
- validPluginName 提取为共享函数,install/remove/read 三处共用,防止检查漂移

禁用是**运行态开关,不删文件**:插件把线上网关搞坏了、但离修好只差一行时,
运维需要把它移出请求路径而不丢失它(同 systemd mask 而非 remove 的道理)。
它**不跨重启保留**——一个悄悄比操作者意图活得更久的"禁用"本身就是个意外。

Builtin 的判定是「加载的源码与内嵌版本逐字节相同」,而不是「名字匹配」:
被改过的 billing.lua 不能被标成 builtin,否则 UI 会提供覆盖用户改动的操作。

on_disk 列表包含**加载失败**的插件。否则一个语法错误的插件在 UI 上直接消失,
运维看到的现象是"插件不见了"而不是"插件报错了"。

## 峰谷 / 时段定价
commandcode 的 DeepSeek V4 系列就是高峰 01-04 & 06-10 UTC 工作日 2 倍价
(非高峰 17h/天)。静态价目表达不了,而算错方向是**静默**的。

价目条目可带 peak = {multiplier, windows=[{days, hours}]}。命中任一窗口即乘。
★ 用 `os.date("!%H")` 取 **UTC** 小时:provider 费率表按 UTC 标注,而网关跑在
本地时区(本机 Asia/Hong_Kong)。混用本地小时会让峰谷整体偏移 8 小时,
白天算成夜间——比不做峰谷还糟。

## ★ 实现与注释不一致,被判据抓住
applyPeak 最初直接 `price.prompt = price.prompt * m`,注释写「缓存读不翻倍」。
但 costFor 里**缓存读价是从 price.prompt 派生的**,所以原地翻倍会把缓存读
也翻倍——两个折扣被叠在一起,而 provider 从没打算叠。
改成 applyPeak 只**记录**乘数,由 costFor 分段应用:fresh prompt 与 completion
翻倍,cache read 那一项不动。
只靠注释说明意图是不够的:TestBillingPeakDoesNotDoubleCacheRead 立刻红了
(0.006 vs 期望 0.003)。变异回原实现仍是红的。

## 判据(21 个计费测试全绿,新增 5 个峰谷)
  窗口恒命中 ×2 / 窗口永不命中保持静态价 / 星期不匹配不命中
  (这条正是防"用本地时区整体偏移 8 小时")/ 无 peak 规则向后兼容
  / 缓存读不随峰谷翻倍

后端部分:构建/vet/gofmt 干净,8 个包全绿。
2026-10-02 08:33:41 +08:00
cb6df0a3f0 fix(billing): 缓存命中按全价计 + 未定价流量静默记 0
部署前审计计费插件时自己找到的两个真缺陷,都会直接算错钱。

## ★ 缺陷 1:缓存命中按全价计(高估约 10 倍)
costFor 只看 prompt_tokens,不区分其中多少是缓存命中。实测(审计脚本,非推演):
1M prompt token 里 900k 是 cache_hit → **算出 10 USD**,而缓存读通常只要 1/10
价,正确值 ~1.9。agent 流量反复重放长前缀,正是缓存要让它便宜的那类流量,所以
这个偏差恰好落在最高频的流量上。

改为拆分:
    fresh  = prompt_tokens - cache_hit_tokens  → 全价
    cached = cache_hit_tokens                  → 全价 × cache_discount
cache_discount 默认 0.1(DeepSeek/Qwen/Kimi 的量级),可按条目覆盖——**折扣率是
每个 provider 的事实、不是自然常数**,所以 0.1 只是默认值而不是硬编码常量。
另外把 cache_hit 钳到 prompt 以内:适配器报出比 prompt 还大的缓存命中数时,
fresh 会变负数,凭空产生负计费 token。

## ★ 缺陷 2:未定价模型静默记 0(最危险)
没有任何价目覆盖的请求,成本记 0,而 **requests 和 token 数照常计入 total**。
于是账单看起来完全正常,只是 quietly 少报——没有任何报错,没有任何异常。
比多算危险得多:多算你会去查,少算你不会知道。

新增两个维度把这件事变成显式信号:
    unpriced_reqs    未定价请求数
    unpriced_models  按模型点名,直接告诉你价目表缺哪一行
仪表盘加一张 "Unpriced" 卡片,**这个数应该是 0**。
任何维度(source / model / key)覆盖了就算 priced。

## 修这两个时自己踩的坑
第一版把未定价统计块写在了 `local s = plugin.state` **之前十行**,
在一个全新插件上 hook 直接抛 "attempt to index global 's'",于是
**整条请求什么都没记**——计费插件能有的最坏失败方式。
是 TestBillingZeroPricesIsSafe 的 "requests = 0" 抓到的。
代价:一个计费插件静默失效,而网关日志里只有一行 hook error。

## 判据(351 个测试全绿,计费相关 16 个)
新增 5 个,全部是**具体金额**断言:
  TestBillingCacheHitsAreDiscounted          1M/900k 命中 → 1.9
  TestBillingCacheDiscountIsPerModel         覆盖为 0 / 1 两种极端
  TestBillingCacheHitClampedToPrompt         荒谬的命中数不产生负费用
  TestBillingCountsUnpricedTraffic           只数未定价的那个,且流量仍计入 total
  TestBillingAnyDimensionCountsAsPriced      源维度定价也算 priced
2026-10-02 01:14:33 +08:00
42764bc99e feat(plugin): AUTO 调度轨迹可见(chain_step stage)
被问"还有 auto 调度相关 stage 呢?"问出来的真实缺口。

## 问题
chainDrive 只返回 (resp, src, model, err),调用方只知道**最终哪个槽位赢了**。
遍历过程中算出来又丢掉的东西——哪些档被跳过、为什么跳过、哪些槽位硬失败、
哪档全忙——一律不可见。ChainErr 里其实有这些,但**只在全部失败时**才填,
而它是 error 返回值不是记录。于是:

    "tier 1 冷却所以降级到 tier 3"  ==  "tier 1 正常接单"

对插件而言 tier 只是个常量 -2("resolved by the chain"),信息量为零。而这
恰恰是优先级链存在的全部理由,也是"我那个贵模型为什么没被用"的答案。

## 做法(scheduler 侧零新依赖)
新增 TraceEvent / TraceSink,chainDrive 多一个可选 sink 参数:

  - TraceEvent 是本包的普通 struct,sink 是 func 参数 ⇒ **不新增 import**,
    scheduler 仍然可独立测试
  - sink 为 nil 时每次 emit 只多一次 nil 判断;没有插件的网关在 AUTO 热路径上
    零开销(gateway 的 chainTraceSink 直接返回 nil)
  - 事件是纯观测:scheduler 不基于它做任何分支,gateway 也不把它喂回路由/
    冷却/配额

四种 kind:tier_skip / slot_fail / tier_busy / selected,selected 每次成功
遍历恰好一次且是最后一步。顺序保证所有 step 在 routed 之前。

## 暴露给插件
新增 chain_step stage(逐个步骤),并在 request_end 载荷里加三个便于做报表的
字段:chain_walk(上限 12 步,防审计记录膨胀)、degraded、tier_served。

## ★ 计费口径(我按推荐的做,已写进文档,需要你确认)
**按实际服务的模型计费**:降级到 tier 3 仍按 tier 3 的价算,轨迹只作观测。
理由与 §7.5 的边界一致——插件只报表不执法,两套口径混在一起会引出"降级该不该
多收钱"这种无法从代码判断的争议。若要改成"按本该用的档计价",需要在 models
价目里允许按 tier 定价,这我没做,因为那是个产品决策。

## 计费插件同步消费
by_tier_served / skip_reasons / degraded_reqs 三个新维度。skip_reasons 的等待
时长做了归一(`no free slot within <wait>`),否则 busy-wait 文案一变就多一行。
降级次数在 request_end 里计而不是在 chain_step 里计:一次降级的请求要走多步,
按步计会重复计数。

## 判据(346 个测试全绿,新增 15 个)
  scheduler  6 个:正常路径只发一个 selected / 跳档+降级可见 / 硬失败与跳档
                严格区分(不可混为一谈,否则抖动上游看起来像空闲上游)/
                nil sink 安全 / 全失败时轨迹与 ChainErr 并存且不互相破坏 /
                空链不发事件
  gateway    1 个端到端:tier 1 全 500 → 插件收到 slot_fail(tier 1) +
                selected(tier 2),request_end 的 tier_served=2 且 degraded=true
  lua        2 个:降级计数与按实际模型计价 / 跳过原因归一聚合
  lua        1 个:chain_step 是真 stage 且顺序正确

3 个变异都红:去掉 slot_fail(3 个判据红)/ 去掉 tier_skip(1 个)/
去掉 degraded 字段(1 个)。
2026-10-02 01:03:39 +08:00
a51a6811a6 feat(plugin): Lua 插件机制 + 计费插件 + 插件文档
插件 = plugin_dir 下的单个 .lua 文件,做两件事:挂请求流水线的钩子、在启动时
贡献 WebUI 界面(整页或往现有页面追加组件)。两者独立。

## 流水线 stage(三个)
  request_start  已解析鉴权、未选源
  routed         已选定 (source, model)、未发往上游
  request_end    每请求恰好一次,带最终计量
request_end 挂在 gateway.writeRec——四条入口路径(直连/AUTO × 流式/非流式)的
唯一汇合点:既不漏(流式 token 只有流结束才知道)也不重。

## 计费插件(plugins/billing.lua,默认 seed,开箱可用)
源 / 模型 / 密钥三个维度定价。token 价优先级 keys > models > default;per_request
固定价是**叠加**的(生图模型可以既算 token 又收固定费)。单位是 USD/单 token,
即各家 provider 的公布口径。累计 total / by_source / by_model / by_key / by_day。
失败请求保留 token 费用、丢弃固定费(可经 count_failures 翻转)。
界面 = 一个独立页 + 状态页顶部一块总开销 tile。

## 一个明确的设计边界
计费插件**只报表,不执法**。网关自己的配额会计(stats.go,入口强制)才是限额
权威,插件不参与任何路由/配额决策。两套独立会计若对不上,比一套功能略少的
更糟。

## ★ 中途改掉的一个根本设计错误
最初让插件复用适配器的**弹性 worker 池**(多状态)。这对适配器是对的(它们无
状态),对插件是错的:计费插件往 plugin.state 累加,多状态意味着总量被劈成
几份;而 SetState 写价格只写进其中一个 worker,钩子恰好跑到另一个时**所有请求
按 0 计费**。改为**单状态 + 互斥锁**。代价写进文档:钩子必须短、同步、不阻塞,
卡住的钩子会卡住所有插件的钩子。
这个 bug 是测试逼出来的——先写了 SetState+Fire 的用例,数字全是 0 才挖出来。

另一个连带缺陷:只带 prices 的 PUT 会整体替换 state,把累计量清零。改为
prices/state 分离——prices 是配置、state 是历史,改价不动账。

## 撞到的三个 Lua 绑定的坑(都写进注释)
  - SetGlobal **会 pop 栈**:连着调两次,第二次从空栈取,赋成 nil
  - GetField 索引越界是 **SIGABRT 整个进程**,不是 panic,recover 救不了
  - Call(nargs, n) **不接受函数索引**,它调的是 nargs 个参数正下方那个;
    传索引会调到参数上("attempt to call a table value")
另外 GetField/SetField 用绝对索引,SetTop(0) 之后必须重取。

## 错误隔离
钩子 error() 不影响转发:捕获 → 记进 hook_errors → 跳下一个插件。适配器出错
会让源进冷却,插件出错**零惩罚**——插件是可选功能。/api/plugins 的 hook_errors
让"坏掉的插件"可见而不是静默消失。

## 界面注入
GET /api/ui-inject 一次返回所有插件的扩展(侧栏需要全部 page 才能建好)。
WebUI 在首次 render **之前** await 注入:先插 HTML 再重建 <script> 让它执行
(innerHTML/template 插入的 script 不会执行,这正是要的效果——避免脚本跑在
自己 DOM 之前)。注入失败不影响仪表盘。
browser 侧 pluginAPI 暴露 fetchState / postState / onTabShown。

## 文档
docs/plugins.md —— 快速上手、加载与热更新、三个 stage 的完整字段表、界面扩展、
状态与 HTTP API、运行时约束(单状态/异常隔离/内置函数)、计费插件的定价与
计费策略、排错表、与适配器的对比表。

## 判据(328 个测试全绿,插件相关 33 个)
  - 计费断言的是**具体金额**(0.00625 / 0.0402 / 0.0075…),不是"能加载"
  - 4 个变异都红:钩子异常不隔离 / prices 清空累计 / 忽略 key 优先级 /
    毫秒时间戳不换算
  - UI 侧 6 个判据把注入顺序、script 执行时机、pluginAPI 名称、tab 路由、
    anchor 四种形式、失败非致命全钉住
  - 鉴权:state 读任意角色、写仅 admin
2026-10-02 00:37:29 +08:00
a2e1adc2d8 fix(gemini): endpoint 自相矛盾导致预置模板必失败
gemini.lua 里 adapter.endpoint = "/v1/models",而它自己的注释写的是
  POST /v1/models/{model}:generateContent
两者矛盾,而 Go 侧是静态拼接(provider.URL = base_url + endpoint),拼不出
模型名。预置模板 "Google Gemini"(base_url=.../v1beta)于是会 POST 到
  https://generativelanguage.googleapis.com/v1beta/v1/models
既多一段 /v1,又缺 :generateContent——那是 Gemini 的模型**列表**端点,
对 POST 返 405。所以任何用户从模板建这个源,拿到的都是必定失败的源。

实测确认影响范围:线上 21 个源里没有 gemini(openai×15 / trae / sensenova /
opencodezen / deepseek / anthropic / agentrouter),所以是潜伏缺陷。

修法:endpoint 改成模板 `/v1beta/models/{model}:generateContent`,新增
provider.ChatURL(model, stream):
  - 用 **PathEscape** 替换 {model}——模型 id 进的是 URL 路径,不转义的话
    一个 "/" 就会静默指向另一个资源(判据里用 RequestURI 而非 URL.Path
    断言,因为后者是解码后的,看不出 %2F);
  - 流式把 ":generateContent" 换成 ":streamGenerateContent"(同一个路径、
    不同动词,也在路径里)。替换刻意只认这个精确后缀,免得别的适配器
    仅仅提到这个词就被改写;
  - source 自己设的 endpoint: 仍然优先,模板被整体跳过。
Chat / ChatStream / probeChat 三处调用点改为传本次请求真实的 model——AUTO
按槽位把 req.Model 钉死,所以 URL 必须跟随**请求**的模型,用源默认模型会让
多模型源每次都打同一个(还记到别的模型的账上)。

判定静态 endpoint 的其他 10 个适配器零影响(TestNonGeminiEndpointsAreUntouched)。

顺带:Stats 的 mutex 不是可重入的,导出方法自己加锁、*Locked 后缀要求调用
方持锁。持锁调导出方法会死锁——我的探针真卡死过一次(直到 10 分钟超时)。
补上 LOCKING 注释,并加判据把这条规则钉住(含一个 20 秒上限的行为判据,
让未来的重构撞死锁时快速失败而不是拖满整个套件)。
2026-10-01 23:37:17 +08:00
97bb9c6ef4 feat(opencode): 透传 completion_tokens_details.reasoning_tokens 与上游 cost
回答「opencodego 的用量与费用透传呢」时逐字段核对上游产出,发现 usage 漏了
一项、费用整项丢失。

## 上游实际发什么(实测 opencode.ai/zen/go/v1)

  {
    "choices": [...],
    "usage": { "prompt_tokens": 37, "completion_tokens": 40, "total_tokens": 77,
               "prompt_cache_hit_tokens": 0, "prompt_cache_miss_tokens": 37,
               "prompt_tokens_details": {"cached_tokens": 0},
               "completion_tokens_details": {"reasoning_tokens": 40} },
    "cost": "0"
  }

cost 在**顶层**且是**字符串**。流式时还会单独发一帧:
{"choices":[],"cost":"0"}

## 此前丢了两样

1. completion_tokens_details.reasoning_tokens —— 输出里有多少是思考 token。
   没有它,客户端无法判断 completion_tokens 里多少是可见回答、多少是思考,
   而两者都按输出计费。
2. cost —— 唯一的费用信号,网关整个丢弃。Go 订阅是包月制恒为 "0",
   但 Zen 按量付费模型(以及未来的其它源)有信息量。

顺带修掉一处流式/非流式不一致:命中缓存时上游同时给
prompt_tokens_details.cached_tokens 和独立的 hit/miss,流式路径写成了 elseif,
只留 details,与非流式产出不同(只认独立字段的老客户端会看不到缓存)。

## 实现

- types.TokenUsage += CompletionTokensDetails;UnifiedResponse / UnifiedChunk += Cost
- opencodego/opencodezen 适配器映射两个字段;空 choices 帧改成 usage 与 cost
  都可带(早退只带 usage 会把同帧的 cost 丢干净 —— 新测试先抓到的就是这个)
- Gateway ChatCompletion / ChatChunk += cost,随终帧发(对齐上游的
  {"choices":[],"cost":"0"} 形态)
- Go 兜底 standardSSEChunk 同步支持(openai 系适配器不再漏 reasoning_tokens;
  纯 cost 帧不再被整体丢弃),新增 rawCostString 兼容字符串/数字两种形态

费用只做**搬运**:不解析、不换算、不汇总 —— 它是上游事实,且只有部分上游提供。

## 验证

经网关实测 gozen:deepseek-v4.1-flash,流式与非流式产出逐字段一致:
  prompt_tokens_details.cached_tokens=6784
  prompt_cache_hit_tokens=6784 / miss=148
  completion_tokens_details.reasoning_tokens=16
  cost="0"

测试:TestOpenCodeCostAndReasoningPassthrough(含「无数据不得凭空造字段」反例)、
TestOpenCodeStreamCacheFieldsMatchNonStream、TestTokenUsageMarshalsCompletionTokensDetails。

(cherry picked from commit c744ee151e)
2026-09-27 17:12:03 +08:00
d81074621a fix(opencode): 采纳客户端真实会话 id + 超窗消息不再被限流措辞封杀
两处都源于同一次排查:pi 到底有没有带会话标识、超窗为什么触发不了压缩。

## 1) 客户端会话 id:pi 一直在发,只是被配置关掉了

之前结论是「通用客户端不发会话 id」——只对了一半。pi 有会话 id,且能发:
pi-ai 的 createClient 在 compat.sendSessionAffinityHeaders 为真时,会把
平台会话 id(uuidv7,整个会话恒定)放到 x-session-affinity /
x-client-request-id / session_id 上。该开关默认 false,而 llmsproxy 的
provider 配置里没开,所以此前一直收不到。

现在网关按优先级采纳:x-session-affinity → x-session-id → session_id →
body 的 prompt_cache_key,并把值经 types.ChatRequest.ClientSession 传到
适配器 meta.client_session。适配器的会号种子优先级变为:
客户端会话 id > 首条 user 消息指纹 > 按源固定。

刻意不采纳 x-client-request-id:名字含 request,部分客户端每请求都换,
拿它当会话会让上游前缀缓存永不命中(pi 总会同时发 x-session-affinity,够用)。

实测:抓 127.0.0.1:8081 的真实 pi 请求,配置打开后收到
x-session-affinity = session_id = x-client-request-id = <子会话 uuid>。
上游缓存确为会话级隔离(同前缀、不同会号:A 冷→命中,B 首次仍为 0),
两个不同 header 值互不命中,反证网关确实采纳了客户端会话 id。

## 2) 超窗消息必须「干净」,否则被同链的限流措辞反向封杀

pi 的 isContextOverflow 先查 NON_OVERFLOW_PATTERNS(/rate limit/、
/too many requests/、Bedrock 前缀),命中就直接判为「非超窗」——**即使
消息里已经有 context_length_exceeded**,pi 也不会压缩重试。

而 AUTO 链的失败消息天生是多 tier 原因的拼接,超窗 tier(gozen 400
maximum context length)常与配额/限流 tier(429 token plan exhausted、
cooling、no free slot)同时出现。此前把 tier 明细原样拼在归一化标记后面,
等于让一条限流 tier 的措辞反过来封杀超窗识别。

现在超窗走独立的干净消息:
  context_length_exceeded: context window is full; reduce the length of
  the messages (gozen/deepseek-v4.1-flash)
只留超窗措辞 + 超窗源名,不带任何其它 tier 的文本。

测试:TestOverflowMessageSurvivesRateLimitedSiblingTier 用 pi 的完整判定
顺序(先 NON_OVERFLOW 后 OVERFLOW)断言同链限流 tier 不再封杀超窗识别;
TestClientSessionFromRequestHeaders / TestClientRequestIDIsNotUsedAsSession /
TestOpenCodePrefersClientSessionID 覆盖会话采纳与优先级。

(cherry picked from commit c9c09b2ba2)
2026-09-27 17:12:03 +08:00
4dd3431c26 feat(opencode): per-conversation session via first-user-message fingerprint
Follow-on to the session-stability fix. "Per source" already made the
prefix cache hit, but it puts every conversation into one upstream session.

Using the client's own session id is not possible: capturing real agent
traffic (tcpdump on 127.0.0.1:8081) shows generic OpenAI clients send NO
session identifier at all — no user / session_id / conversation_id /
metadata in the body, and no session header (only X-Stainless-* plus
User-Agent: pi). The x-opencode-session the Go endpoint asks for is an
OpenCode native-client concept that a generic client cannot forward.

Since history is replayed every turn, the FIRST user message is invariant
for the life of a conversation, so it is used as the conversation
fingerprint. The session becomes stable within a conversation and distinct
across conversations; requests with no user message fall back to per-source
stability.

Measured through the gateway (same 5.7k-token prompt): 2nd call
cached_tokens=5504, and an unrelated conversation gets its own session.

Test: TestOpenCodeSessionIsStableForCache covers same-conversation
stability, cross-conversation separation, per-request request ids and the
sessionless fallback.

(cherry picked from commit 39b48e556e)
2026-09-27 17:12:03 +08:00
3d1407e6fa fix(opencode): make x-opencode-session stable so the upstream prefix cache can hit
The opencode adapters derived x-opencode-session from meta.timestamp, i.e. a
brand new session on every request. The upstream prefix cache is
session-scoped, so no request could ever hit it, and the cache fields the
endpoint does report (prompt_tokens_details.cached_tokens,
prompt_cache_hit_tokens/prompt_cache_miss_tokens) always came back 0/absent.

Measured against the live endpoint, same 6032-token prompt:

  fixed session id   -> 2nd call: hit 5888, miss 144
  rotating session id -> every call: hit 0, miss 6032

Fix: derive the session from the source name (stable), matching how
x-opencode-project is already derived. x-opencode-request stays unique per
request — it is only a request identifier, not part of the cache key.
Applied to both opencodego and opencodezen.

Through the gateway the same prompt now reports, on the 2nd call:
  details={'cached_tokens': 5888} hit=5888 miss=144      (non-streaming)
  prompt_tokens_details={'cached_tokens': 5888}          (streaming)

Test: TestOpenCodeSessionIsStableForCache asserts the session is stable
across requests for one source while the request id differs.

(cherry picked from commit 791d198f47)
2026-09-27 17:12:03 +08:00
db320d2f6c fix(opencodego): inject empty reasoning_content on tool-calling turns
v1.5.4 stopped stripping reasoning_content, which fixes clients that send
it — but most agent clients (pi included) never store or replay their
reasoning, keeping only the tool call. OpenCode Go validates the field on
any assistant turn that carries tool_calls and rejects the whole request:

  400 invalid_request_error: The `reasoning_content` in the thinking mode
  must be passed back to the API.

Verified against the live endpoint that an EMPTY string satisfies the
check, so the adapter now fills in "" when a tool-calling assistant turn
has no reasoning_content. Nothing is fabricated: the reasoning shown to
the client is still exactly what the upstream returned for that turn.

Measured: with a tool_call + tool_result history and no reasoning_content,
all 25 configured Go models returned 400 before and all 25 answer
correctly now.

Test: TestOpenCodeGoVsZenReasoning also pins that a plain assistant turn
(no tool calls) must NOT gain the field.

(cherry picked from commit d1a72cd23a)
2026-09-27 17:12:03 +08:00
b544671f5e feat(adapters): split opencode into opencodezen and opencodego
Zen (https://opencode.ai/zen/v1) and Go (https://opencode.ai/zen/go/v1)
are different services with different requirements, and one shared adapter
could not satisfy both.

The decisive difference is reasoning_content:

  * OpenCode Go runs thinking models and REQUIRES the assistant turn's
    reasoning_content to be echoed back. The shared adapter stripped it
    (msg.reasoning_content = nil), so every replay of a thinking turn
    failed with:
      400 invalid_request_error: The `reasoning_content` in the thinking
      mode must be passed back to the API.
    Reproduced directly: the same request with reasoning_content -> 200,
    without -> 400. That is why the Go tier never worked in an agent loop.

  * The Zen free pool must not receive it, so it keeps stripping.

Both adapters keep the earlier fixes they share (never drop an assistant
turn carrying tool_calls; send stream_options only when streaming; role
whitelist; multimodal strip) and the opencode client fingerprint headers —
the Go endpoint additionally REQUIRES x-opencode-session, which the
adapter already sends.

config: localzen -> opencodezen, gozen -> opencodego.
Verified: all 25 gozen models answer correctly through the gateway with a
thinking + tool_call + tool_result history (was 0/25 before), streaming
included; the Zen free models still pass.

Test: TestOpenCodeGoVsZenReasoning pins the Go-keeps / Zen-strips split.
(cherry picked from commit 71e9040a45)
2026-09-27 17:12:03 +08:00
0881a49286 fix(opencode): only send stream_options with stream:true
OpenCode Go (and other strict OpenAI-compatible upstreams) reject a
non-streaming request that carries stream_options with
"stream_options should be set along with stream". The adapter attached it
unconditionally, so every non-stream call through the opencode adapter
failed on those upstreams.

Verified against OpenCode Go: 25/25 configured models now pass a real
completion through the gateway (they previously 400'd).

Test: TestOpenCodeStreamOptionsOnlyWhenStreaming (absent when
non-streaming, include_usage present when streaming).

(cherry picked from commit 0528941e24)
2026-09-27 17:12:03 +08:00
4e2b5bfda9 fix: empty-array content becomes invalid {} on every pass-through adapter; gemini/ollama drop tool calls
Three related forwarding defects found by auditing every adapter with a
tool-calling replay (assistant turn with content:[] + tool_calls).

1) content:[] -> content:{} (all 12 openai-adapter sources, plus
   deepseek/trae/sensenova/agentrouter/github/groq/kimicode/mistral)

   Lua adapters json.decode the request and re-encode it, and an empty Lua
   table is indistinguishable from an empty JSON array — the encoder emits
   {} for both. Agent clients serialise a tool-calling assistant turn with
   no text as content:[], so every pass-through adapter rewrote it to
   content:{} — not valid OpenAI (content is string|array|null). Verified
   against a live upstream: content:[] produced "400 invalid arguments"
   while content:"" was accepted.

   Fixed once at the decode boundary (types.ChatMessage.UnmarshalJSON):
   empty-array content normalises to "" and an empty tool_calls array is
   dropped, so every adapter — including future ones — sees a valid shape.

2) gemini dropped tool_calls and never emitted functionCall /
   functionResponse; the tool role also stayed as an invalid role inside
   contents and system was not moved to systemInstruction.

3) ollama copied only role/content, dropping tool_calls and the call
   attribution entirely (it needs tool_name, not tool_call_id).

Test: TestAdaptersPreserveToolCalls asserts, for every adapter, that the
call id (or function name where the wire format has no id), the function
name, the tool result and the trailing user turn all survive, plus a
negative control for plain text.

(cherry picked from commit 9114468753)
2026-09-27 17:12:03 +08:00
61b847da86 fix(opencode): never drop an assistant turn that carries tool_calls
An agent client (pi) serialises an assistant turn whose content is only
[thinking, toolCall] as content:[] with tool_calls. The multimodal-strip
pass treated an empty content array as 'nothing left, drop the whole
message' and discarded the tool_calls with it.

The next message is that call's tool result, so it arrived orphaned: the
model saw a result for a call it had never made and re-issued the same
call on every turn — an endless repeated-tool-call loop. Reproduced
against a capture sink: content:[] lost tool_calls, while content:"" and
content:null kept them.

Only messages with neither usable content NOR a tool call now get
dropped. Content that collapses to empty but still has tool_calls or a
tool_call_id is emitted as "" instead.

Test: TestOpenCodeKeepsToolCallWithEmptyContent (plus a negative control
that an image-only message without tool calls is still dropped).

(cherry picked from commit 499f0cac2f)
2026-09-27 17:12:03 +08:00
6432b61eb6 fix(gateway): AUTO scope grants all models — restrict to routing mode only
A key with scope=[AUTO] could previously:
1. request ANY concrete model id directly (hasScopeModel/checkModelScope
   treated AUTO as a wildcard)
2. see the full 56-model list on /v1/models (intersectModels considered
   AUTO as grant-everything)

AUTO now only authorizes the AUTO routing mode. Direct requests to a
specific model require an explicit scope entry.

Also carries agentrouter.lua WAF fingerprint headers (Origin/Referer/
X-Requested-With) already staged on this branch.

Tests: TestHasScopeModelWithSourcePrefix updated; full suite green.
(cherry picked from commit 7fb8f96b82)
2026-09-27 17:12:03 +08:00
c6c3e0dcd7 fix(agentrouter): sanitize tool-call ids — it fronts Claude too
Full-repo audit after the anthropic/openai fix: agentrouter exposes
claude-opus-4-8, so it inherits Anthropic's tool id rule
^[a-zA-Z0-9_-]{1,64}$ and rejects the whole request on a violation, exactly
like justwoker/tabitoken/扇贝. It was the only remaining adapter serving Claude
models without the sanitizer, so a client that had picked up a dirty id (e.g.
"bash:0" from moonshotai/kimi-k3) would still lose every turn here.

Same shape as the other two — inbound tool_calls[].id + tool_call_id, outbound
non-streaming ids and the first streamed fragment — and the test now asserts
all three adapters rewrite an identical input identically, so a client mixing
sources within one session cannot end up with unpaired tool calls.

Audit result: every source exposing a claude/opus/sonnet model (qijiar, toter,
juziai, agentrouter, justwoker, api456, tabitoken) now routes through a
sanitizing adapter.
2026-09-06 10:08:52 +08:00
6bdb9fcc44 fix(anthropic): count cached input in prompt_tokens instead of dropping it
Anthropic and OpenAI disagree on what the prompt count means:

  Anthropic: input_tokens EXCLUDES cached blocks; cache_read_input_tokens and
             cache_creation_input_tokens are separate, additive, billed input.
  OpenAI:    prompt_tokens INCLUDES its cached_tokens subset.

anthropic.lua mapped input_tokens straight onto prompt, so a cache-heavy turn
was doubly wrong: the billed prompt was undercounted by the entire cache
portion, and cached_tokens could exceed prompt_tokens — a cache hit rate above
100% for any client that divides one by the other. cache_creation_input_tokens
was never read at all, so a cache-write turn silently lost those billed tokens.

Worse, the streaming path dropped the cache split entirely: message_delta
carries the FINAL usage and only mapped input/output, so every streamed
response reported no cache information even when the upstream sent it.

All three counts are now summed into prompt, with the read half exposed as
prompt_tokens_details.cached_tokens plus the DeepSeek-legacy hit/miss pair, via
one shared map_usage() used by transform_response, message_start and
message_delta. A reported zero stays distinguishable from "never reported": the
split is emitted whenever either cache field is present, and omitted entirely
when the upstream mentions neither (justwoker reports only input/output plus its
own cost fields, so its output is byte-identical to before). map_usage returns
nil for a countless object, preserving "no usage in this chunk means say
nothing" rather than reporting zeros.

message_start's placeholder count is still emitted: justwoker reports 160 there
and the real 6931 in message_delta, and the gateway's mergeUsage lets the later
non-zero value win.
2026-09-06 09:55:22 +08:00
131a42a169 fix(adapters): sanitize tool-call ids so one bad upstream can't kill every Claude slot
Anthropic requires tool_use.id / tool_result.tool_use_id to match
^[a-zA-Z0-9_-]{1,64}$ and rejects the WHOLE request otherwise with
REQUEST_BODY_INVALID / "Invalid tool use format". OpenAI has no such rule, so
an OpenAI-compatible model can mint an id like "bash:0"
(xinjianya/moonshotai/kimi-k3 does exactly that).

In a fan-out router that id does not stay local: the client stores it in its
history and replays it to every other source. One such id therefore kills
every Claude slot at once — justwoker, tabitoken and 扇贝 are all
Claude-behind-{OpenAI,Anthropic} — and an AUTO request falls through all four
tiers to whatever tolerant model is left. Observed live: 4 consecutive 503
"all N auto providers failed" with tier 1/2/3 each reporting the same 400.

Both directions are sanitized, in both adapters:
  - request:  tool_calls[].id and tool_call_id, so poisoned history recovers
  - response: non-streaming tool_calls[].id and the first streamed fragment,
              so a bad id never enters a client session in the first place

safe_tool_id is pure and deterministic, so a call and its result are rewritten
identically within one request. A rewritten id keeps an 8-hex digest of the
original, without which distinct ids could collapse ("a:b" and "a_b") into a
duplicate/unpaired tool_use. Already-legal ids pass through byte-identical, so
well-behaved traffic is unaffected. openai.lua carries its own copy because
Lua adapters have no shared prelude.

Streamed argument fragments carry no id and must stay id-less, otherwise
index-based accumulation on the client breaks; a test pins that.
2026-09-05 22:18:31 +08:00
4d197e4bd3 fix(trae): recover legacy [Called tool:...] tool call format
Root cause: trae-local-api is deployed to users without the "Fold past
assistant tool_calls into [Called tool: name({...})]" text format that
their own histories already contained. Trae-Local-API-LLM then mimics this
format in subsequent responses. The trae.lua parser only recognized
<tool_call>...</tool_call> or <toolcall>...</toolcall> tags, so
[Called tool: ...] responses were left unparsed and the client received
plain text where a structured tool_calls array should be.

Fix: Add a legacy pattern match at the END of parse_text_tool_calls to
catch the [Called tool: name({args})] shape and emit proper tool_calls.
This is a fallback; models should emit <tool_call> tags per system prompt,
but we tolerate the mimicked form for robustness.
2026-09-05 09:46:36 +08:00
2e3d5b79ad fix(adapters): stop dropping non-streaming tool calls (agent loops died on turn 2)
Four adapters handled tool_calls in transform_stream_chunk but lost them in
transform_response, so any NON-streaming tool-using conversation broke on its
second request: the client received finish_reason:"tool_calls" with no
tool_calls payload, replayed an assistant message whose function
name/arguments were empty, and the upstream rejected the next turn with

    400 invalid tool_call function, function/name/arguments cannot be empty

The production audit trail shows 46 such failures on sensenova alone.

- sensenova.lua: forward message.tool_calls, decoding the arguments JSON string
  into an object as the unified shape expects.
- gemini.lua: collect functionCall parts from candidates[].content.parts. Also
  correct finish_reason, since Gemini reports "STOP" even when it emitted a
  function call and clients keyed on it treat that as a finished answer.
- ollama.lua: the field was initialized to an empty table and never filled;
  fill it and likewise correct done_reason "stop" -> "tool_calls".

trae is a different failure with the same symptom: trae-local-api's OpenAI
endpoint (/v1/chat/completions, src/server.js:353) never reads the request's
`tools` array — only its Anthropic endpoint does — so the relayed model is never
told the tool schema and instead PRINTS a <tool_call>{...}</tool_call> block into
content, leaving message.tool_calls null and finish_reason "stop". An OpenAI
client sees an ordinary completion and its agent loop ends mid-conversation.
trae.lua now recovers the structured call from that text, strips the block from
user-visible content, and corrects finish_reason. Both tag spellings
(<tool_call>/<toolcall>, the latter is what the same codebase's Anthropic prompt
asks for) and all three argument key names (arguments/params/input) are accepted.
This is a defensive fallback: fixing the upstream shim to honour `tools` remains
the real fix, since the model still guesses parameter names.

Tests: TestNonStreamToolCallsPreserved covers all ten OpenAI-shaped adapters,
TestGeminiNonStreamToolCalls and TestOllamaNonStreamToolCalls cover their native
shapes, TestTraeTextToolCallRecovery covers both tag spellings, prose around the
block, and asserts a plain text answer never gains tool_calls.

Verified end-to-end against mock upstreams reproducing each shape: a full
two-round agent loop (tool call -> tool result -> final answer) now completes for
both the structured and the text-emitted variants.
2026-08-31 10:23:35 +08:00
813de19bd0 fix(lua): bound prewarm by queue depth, not by the ceiling
Production oscillated between ~46 MB and ~56 MB RSS with the openai pool
cycling 1 -> 10..12 -> 1 states every couple of minutes, while peak_in_use
never went above 2.

Cause: the batch prewarm sized itself purely on the adapter's ceiling. With
max_concurrent summing to 76, growStep is 8, so any two overlapping requests
warmed 8 states — 6 more than anything was waiting for. A minute later the
janitor correctly reclaimed the surplus, the next pair of overlapping requests
warmed 8 again, and the pool churned boot/discard forever. The elasticity was
working; the growth signal was simply wrong.

Prewarm is now bounded by BOTH limits: the ceiling still caps the step, but the
batch never exceeds p.waiting, the number of goroutines actually blocked on the
pool. Overlapping-but-not-queued traffic (the common case) creates exactly the
states it uses; a genuinely queued burst still ramps in one jump.

TestContentionBatchPrewarms is rewritten to queue real waiters instead of
relying on the ceiling to imply demand, and TestNoPrewarmWithoutWaiters pins the
production shape: two overlapping requests against a 76-wide adapter must create
exactly 2 states.
2026-08-30 08:25:18 +08:00
25d8bd8632 fix(lua): reclaim burst leftovers while traffic continues
Production after the first deploy showed the openai pool stuck at 9 states with
peak_in_use=1: a startup burst grew it, and then it never shrank again. The
shrink grace counter was reset by every CHECKOUT, so on a gateway that always
has a request in flight the counter never reached shrinkGraceRounds and the
burst's leftover states were pinned indefinitely — the same monotonic-growth
behaviour this series set out to remove, just with an extra step.

Grace is now reset by GROWTH (a miss that had to boot a state), which is the
actual signal that capacity is short. Ordinary sequential traffic no longer
defers reclaim, while two consecutive quiet-ish rounds are still required so a
gap between two bursts does not tear the pool down.

TestCheckoutResetsGrace asserted the old behaviour and is replaced by
TestGrowthResetsGrace (sequential traffic must NOT defer, growth must) plus
TestBurstLeftoverIsReclaimedUnderSteadyTraffic, which reproduces the production
shape: 9 concurrent holds, then one request per janitor round, and the pool must
still fall back to the resident floor.
2026-08-30 08:13:34 +08:00
02cb8c0e33 fix(adapters): reflow live sensenova reasoning field, add trae adapter
sensenova: the deployed /etc/llmsproxy/adapters/sensenova.lua carried a fix that
never made it back into the repo — sensenova-6.8-flash-lite reports its chain of
thought in `reasoning` rather than `reasoning_content`, in both the single-shot
response and the stream deltas. Without this the model's output looked empty.
Repo and deployment now match byte for byte.

trae: the adapter was in use on this deployment but untracked, so a fresh
install had no way to serve the trae source.
2026-08-30 08:06:22 +08:00
df9aeed5fb feat(lua): elastic adapter worker pools instead of monotonic growth
ConfigureConcurrency() set each adapter pool's target to the sum of
max_concurrent over its sources (108 on this deployment) and `created` only
ever went UP: once a Lua state was booted it was parked forever, so a
long-running gateway's resident state count was a high-water mark of all
traffic it had ever seen, never of what it currently needs.

Pools are now sized from three live inputs:

  * the adapter's MAX CONCURRENCY (sum of max_concurrent) is a ceiling, not a
    preallocation, and it sets the growth step:
    growStep = clamp(ceil(maxW/8), 1, 8). A 64-wide adapter warms 8 states at
    once on a spike; an 8-wide one creeps up one at a time.
  * the LIVE CONNECTION COUNT (inUse, i.e. checked-out states) sets the shrink
    step: shrinkStep = clamp(ceil(excess/(1+inUse)), 1, excess). With no
    connections the slack collapses in a single round; a busy adapter gives up
    one state per round so the hot path keeps its warm states.
  * how many states already exist (created / len(idle)) decides how much room
    is left to grow and how much can be reclaimed.

Batch prewarm only fires on genuine contention (a miss while every existing
state is checked out), so a single sequential caller keeps reusing one state
rather than burning a whole grow step on a cold start. A single VM-level
janitor goroutine (not one per adapter) reclaims idle states every 30s, and
shrinkGraceRounds=2 plus idleHeadroom=1 keep a gap between requests from being
mistaken for the end of a load period; any checkout resets the grace counter.

`created` now decrements on reclaim and on shutdown, and release() closes a
state outright when the ceiling was lowered underneath it, so shrinking
max_concurrent in the config gives memory back immediately instead of parking
orphans until restart.

PoolStats() exposes created/idle/in_use/waiting/max/resident/grow_step/
shrink_step/peak_in_use for the status API.

Measured on the test instance (single mock source, max_concurrent=64):
idle created=1; 50 concurrent requests -> created=10 (peak_in_use=5, ceiling
respected); after 95s of silence -> created=1. On production after deploy: 13
adapters, 1 resident Lua state total with ceilings up to 76.
2026-08-30 08:05:13 +08:00
624fd74b45 fix: anthropic tool-call round-trip, cache zero-hit parity, round-robin load balancing
anthropic.lua v3.0.0:
- Issue 1: tool_result/tool_use round-trip
- Issue 3: thinking default OFF (opt-in via extra_body.thinking)
- Issue 4: tool_choice mapping
- Issue 5: collect_blocks preserves unknown part types
- message_stop no longer emits done=true (was overwriting tool_calls finish_reason)
- cache_read_input_tokens normalized even at 0

gemini.lua:
- transform_response was missing cachedContentTokenCount

openai.lua (Issue 6):
- transform_error handles flat envelopes, nginx HTML, bare text

chat.go mergeUsage:
- Keep PromptTokensDetails even when CachedTokens=0

scheduler.go:
- Remove sort.SliceStable by Pref; round-robin cursor is the only LB mechanism

provider.go ModelAvailable:
- Also check Pref() > prefMin, persistently failing slots exit cands

presets.go:
- 17 built-in source templates

Tests: 6 new test functions, 2 updated for new semantics
2026-08-28 12:02:46 +08:00
dev
21ec8f59d8 feat: surface zero cache hits — distinguish 'missed' from 'not reported'
Live testing across the zen pool showed models report
prompt_tokens_details.cached_tokens even when the hit count is 0 (e.g.
nemotron-3-ultra-free returns cached_tokens:0, audio_tokens:0,
cache_write_tokens:0). The previous >0 guard dropped those objects, so a
cache-enabled upstream looked identical to one without cache support.

- types: PromptTokensDetails.CachedTokens always emitted (drop inner
  omitempty) so clients see cached_tokens:0 explicitly; dsh reads it as
  a 0% hit instead of 'no data'
- adapters (9): forward prompt_tokens_details whenever the upstream
  provides it (presence check instead of >0)
- Req: add cache_reported flag set when usage carried cache accounting;
  WebUI shows an amber 0% tag for reported-but-missed rows and keeps
  the em-dash only for sources that never report cache data
2026-08-25 10:04:44 +08:00
dev
ec89daad62 fix(adapters): pass through cache tokens in the remaining 8 adapters
Live testing proved both sensenova and zen DO return cache fields:
- zen laguna-s-2.1-free: usage.prompt_tokens_details.cached_tokens = 32
  (real hit), plus cache_write_tokens/audio_tokens
- sensenova glm-5.2: prompt_tokens_details.cached_tokens present (0 on
  short prompts)

The previous round only patched deepseek/openai/anthropic/gemini.lua;
sensenova/opencode (localzen!) and the other adapters still dropped them.

- sensenova/opencode/groq/mistral/github/kimicode: stream + response
  cache passthrough (same pattern as openai.lua)
- agentrouter: response passthrough + NEW stream usage forwarding (it
  previously dropped the terminal usage-only chunk entirely)
- ollama skipped intentionally: its native API has no cache fields

Verified end-to-end through the gateway: localzen/laguna-s-2.1-free now
returns prompt_tokens_details.cached_tokens=32 to clients, and the request
record carries cache_hit_tokens (both chat and stream paths).
2026-08-25 09:52:55 +08:00
dev
045ecf47bc feat(types): pass through upstream cache tokens in TokenUsage
dsh displays cache-hit %, but llmsproxy dropped every upstream's cache
fields — deepseek prompt_cache_hit_tokens, OpenAI prompt_tokens_details.
cached_tokens, anthropic cache_read_input_tokens, gemini cachedContentTokenCount.

Changes:
- TokenUsage: add PromptTokensDetails (with CachedTokens) + PromptCacheHit/Miss
- MarshalJSON: emit prompt_tokens_details.cached_tokens (OpenAI v2 standard)
  and prompt_cache_hit/miss_tokens (DeepSeek legacy) — dsh reads the former
  first, falls back to the latter
- mergeUsage: preserve cache fields across stream chunks
- standardSSEChunk: parse the upstream raw prompt_tokens_details too
- deepseek.lua: forward prompt_cache_hit/miss_tokens + create
  prompt_tokens_details from them
- openai.lua: forward prompt_tokens_details.cached_tokens and legacy
  prompt_cache_hit/miss_tokens; normalize legacy hits into the standard
  object so dsh sees them regardless of upstream format
- anthropic.lua: map cache_read_input_tokens → prompt_tokens_details
- gemini.lua: map cachedContentTokenCount → prompt_tokens_details
2026-08-25 07:57:32 +08:00
b2183df1e8 feat(adapter): move per-source error condensing into transform_error hooks
Every upstream formats errors differently, which is adapter territory:
the protocol gains an optional transform_error(status, body) hook and all
built-in adapters implement their own envelope parsing (zen free-pool
labels, anthropic/gemini/ollama/mistral shapes, sensenova quota notes,
agentrouter WAF pages). The core keeps a single uniform fallback: when no
hook yields a reason clients get "api error <status>: unknown error" and
the raw body goes to server logs only.
2026-08-24 19:17:36 +08:00
bb3af3bdb3 feat(gateway): pass through upstream finish_reason end-to-end
The gateway hardcoded "stop" on every terminating stream chunk, so
tool-call rounds reported finish_reason=stop and length caps were
invisible to clients. UnifiedChunk now carries finish_reason; adapters
emit it (with empty-string finish reasons like sensenova treated as
non-terminal), standardSSEChunk passes it through for un-adapted
upstreams, [DONE] no longer emits a duplicate reason-less done chunk,
and both streaming paths emit the real reason with "stop" as fallback.

Also vendor sensenova/agentrouter adapters into the repo: they were
WebUI-only uploads and a deploy sync silently removed them while live
AUTO-chain slots still referenced them.
2026-08-24 15:05:50 +08:00
9c99ded8c0 fix(adapter): send x-opencode-* identity headers so zen stops returning empty tool-call replies
zen fingerprints clients via UA + x-opencode-client/session/request/project
headers; requests without them are routed as anonymous and fail with a
single-chunk network_error stream (tools+stream) or 503 Endpoint is
unavailable (non-stream), which the adapter laundered into empty-but-valid
replies. Derive per-request session/request ids from build_headers meta
(sandbox has no os/math), bump UA to the real client format, map zen
reasoning field to reasoning_content, and set stream_options.include_usage.
2026-08-24 14:37:09 +08:00
3069cfce4e fix(gateway): pass through upstream token usage in streams for all adapters
The prior usage-passthrough fix only covered openai/opencode; the same
empty-choices+usage drop bug remained in the 5 sibling OpenAI-compatible
adapters, and non-OpenAI providers (anthropic/gemini/ollama) never surfaced
streaming usage at all.

- deepseek/github/groq/kimicode/mistral: preserve usage on empty-choices
  chunks and attach it to normal chunks (same pattern as openai.lua)
- anthropic: emit usage from message_start (prompt) and message_delta
  (completion); gateway merges split usage additively
- gemini: read usageMetadata in the stream path
- ollama: fix non-streaming key (usage -> token_usage, matches
  UnifiedResponse json tag) and read prompt_eval_count/eval_count;
  surface counts from the done stream chunk
- gateway: mergeUsage combines usage across chunks (non-zero fields win,
  total recomputed from prompt+completion) so split usage doesn't lose
  the prompt half; single-chunk case (OpenAI) preserved exactly
- usage-only chunks: done=false (no redundant terminal stop), matching
  the Go fallback standardSSEChunk
2026-08-18 23:24:01 +08:00
83e6d88813 feat(gateway): pass through exact upstream token usage in streams
Streaming responses now carry the upstream's real token usage instead of
gateway estimates:
- UnifiedChunk gains an optional Usage field; adapters (opencode, openai)
  extract usage from upstream stream chunks (including the final chunk with
  empty choices) and pass it through.
- standardSSEChunk preserves usage for passthrough adapters.
- Gateway emits the exact usage in the final stream chunk when available,
  falling back to estimates only when the upstream provided none.

Non-streaming usage was already fixed to emit OpenAI-standard keys.
2026-08-18 19:04:44 +08:00
e356885130 test(lua): add opencode role normalization test 2026-08-16 17:51:43 +08:00
47f3b44d92 feat(gui): Electron desktop with embedded core, tray, autostart, win/linux packaging
- cmd/gui: Electron shell (Clash-Verge style) embedding the full WebUI 1:1
  - embedded llmsproxy core (luajit) with auto-generated profile
  - key stored in keys[] (non-seed) so no replace-the-key warning
  - gw_key cookie injection: web UI works without login
  - side-rail toggles for autostart / silent start
  - system tray with status + controls, silent start (--silent)
  - win cross-build (mingw luajit exe + dll) / deb / AppImage via electron-builder
- Makefile: build / gui / gui-dist / gui-deb / gui-win targets
- README: desktop GUI section
- lua(adapter): opencode normalizes non-whitelisted roles to system
2026-08-16 09:53:05 +08:00
40e08b14d2 fix: zen adapter strips multimodal parts (upstream is text-only) 2026-08-13 22:08:15 +08:00
2bc1d0e67a feat: opencode zen adapter + first-run config generation, fix stats/stream bugs
- adapters/opencode.lua: opencode.ai zen free pool adapter — sends the
  opencode client User-Agent (zen fingerprints clients by UA; non-official
  clients hit FreeUsageLimitError); pairs with api_key: public
- config: no config file ships in the repo; first run generates a default
  config at the -config path with a random admin key, loopback listen and a
  keyless zen source (config.EnsureDefault); remove config.example.yaml
- lua: seed bundled adapters from the embedded FS instead of a hardcoded
  name list
- ui: widen model kind select (chat was clipped to 'cha')
- phase 5 bugfixes: stats ms/s bucket mixing, cleanScopes nil, ctx.Err
  guards, direct-path ModelAvailable, empty stream body failure,
  bestImageModel rewrite, transform failure recording, Core.mu, timer,
  effective model for tool-calls
2026-08-13 12:25:07 +08:00
41c9b0e14a feat: source-model routing disambiguation (source-model/:/ prefix); same-tier round-robin load balancing; public_base_url for copy config; fmtTok(B/M/K) unit scaling; status column reorder (reachability first); drop emoji from seed-warn modal; fix tests for first-run adapter seeding; install lua5.1 dev lib 2026-08-10 12:55:22 +08:00
f46b02089c feat(delete): real deletes — adapters seeded once into adapter_dir (existing dir is authoritative), sources removed from config.yaml on delete; drop tombstone mechanism for both. Docs: multi-key architecture, gateway_keys as seed 2026-08-10 11:52:51 +08:00
5d50b69153 fix: tool call anchor & wire format, streaming chunk passthrough, WebUI narrow-screen, docs bilingual 2026-08-08 11:26:58 +08:00
ccbcede581 feat: LuaJIT worker-pool VM, multimodal/disable-thinking passthrough, WebUI redesign, DeepSeek V4 2026-08-07 13:34:16 +08:00
f7f76e097d feat: ModelRouter — unified OpenAI-compatible multi-source LLM gateway
- Lua adapters per upstream (transform_request/response/stream_chunk, build_headers signing hooks)
- AUTO priority routing with per-model kind (chat/image), explicit source/model routing
- Per-source concurrency caps with queueing, exponential backoff, AUTO failover
- OpenAI-compatible API: chat completions, SSE streaming, image generations, models
- Gateway key auth, web UI for adapter/source management, runtime persistence
- e2e test running the real binary against mocked upstreams
2026-08-05 15:25:47 +08:00