Commit Graph

171 Commits

Author SHA1 Message Date
70f1c879bd fix(startup): 密钥警告改读真实生效的 key 集合
启动时那条「gateway_keys is EMPTY — without a key every request is rejected」
读的是 legacy 的 cfg.GatewayKeys 段,而鉴权实际用 cfg.Keys(core.ListKeys)。
seedKeys 首次启动把 gateway_keys 搬进 keys[] 之后,YAML 里那个列表就不再
被鉴权使用。于是在它被清空(例如轮换掉 starter key 之后)而 keys[] 仍有
7 把可用 key(含 admin)时,进程每次启动都谎报「所有请求都会被拒绝」。

实测:生产日志出现该警告,而同一个 key 请求 /v1/models 返回 200。

- main.go 改为检查 c.ListKeys(),文案改成不绑定字段名。
- 顺带删掉 gateway.New 的 gatewayKeys 参数:函数体从未使用它,
  只读 ListKeys(),留着会继续诱导人以为鉴权来自那个列表。

判据:e2e/TestStartupWarningReflectsRealKeysNotLegacyList —— 构造
「gateway_keys 空 + keys[] 有 key」的真实形态,先断言该 key 确实能鉴权,
再断言日志里不再出现那句谎报。变异验证:回退成 GatewayKeys() 即变红。
2026-09-28 23:42:48 +08:00
04e544c823 chore(version): 1.7.3 -> 1.7.4
发行包不再内置可用 admin key 的修复,走 patch 发布。
v1.7.4
2026-09-28 23:27:54 +08:00
7e33d11d15 fix(packaging): 发行包不再内置可用的 admin key
打包时把本地 config.yaml(gitignored,含运维真实密钥)原样复制成
config.example.yaml,而 postinst 在首次安装且 /etc 无配置时又把它
cp 成生产配置 ⇒ 每次安装都得到一个同值的、公开已知的 admin key。
实测该 key(sk-gw-local-0001)在生产上真实有效(/v1/models 返回 200,
而网关监听 0.0.0.0)。

三处修正:
- 新增 packaging/config.example.yaml(受 git 跟踪的净化模板),
  gateway_keys 留空、sources 留空,并写明不要填死值。
- core-dist.sh / nfpm.yaml 改为打包该模板,不再碰本地 config.yaml。
- postinst.sh 不再投递示例配置:留空文件会让网关启动但拒绝所有请求
  (无门可入)。改为让二进制首启时自行生成随机 admin key 并打印 ——
  每次安装都不同,且开箱可用。示例文件仅作为 /usr/share 下的参考保留。

实测首启:生成 sk-gw-838d66a1... 并打印,与旧的共享固定值不同。
2026-09-28 23:27:42 +08:00
cd82835f25 chore(version): 1.7.2 -> 1.7.3
启动重复播种 admin key 与配置封存非幂等的修复,走 patch 发布。
v1.7.3
2026-09-28 22:53:19 +08:00
26ea782350 fix(core): 修复启动重复播种 admin key + 配置封存非幂等
根因是 unseal 时序:NewFromConfig 把解密放在最后,而之前几步已经在读凭据。

1. seedKeys 重复播种(生产已累积 4 个同名 admin key)
   seedKeys 用 cfg.Keys[i].Key 与明文 gateway_keys 比对去重,但此时内存里的
   key 还是密文 enc:v1:…,比对永不命中 ⇒ 每次重启追加一个同值 admin key。
   实测:core.New(path) 连续重启,seeded key 数 2→3→4 递增。
   (旧测试用 NewFromConfig 构造全新内存对象,没有「盘上已有密文」这个前提,
    复现不出 —— 必须走 core.New 这条读盘的生产路径。)

2. 启动恒重写 config.yaml
   migratePlaintextSecrets 按内存状态判断,而 Save() 末尾会把内存恢复为明文,
   于是每次调用都判定「还有明文」并重写;注释却自称幂等。
   改为 UnsealSecrets 在解密前记录「盘上是否明文」,SealIfNeeded 据此决定
   是否写回 ⇒ 已封存的配置启动不再落盘。

原测试 TestMigratePlaintextSecretsIsIdempotent 用 ModTime 比较,两次写落在同一
时间戳刻度内就看不出来,所以表现为 ~1/6 概率的 flake 而非稳定失败。已改为比较
文件内容并走真实启动路径(UnsealSecrets + SealIfNeeded),并顺带消除该 flake。

附带更正:先前判断「rebuildRegistry 也会拿到密文 API key」不成立 ——
mergedSources → resolveSourceKey 对每个 source 独立解密(belt-and-braces),
provider 始终拿到明文。unseal 前置仍予保留,以消除对该兜底路径的隐性依赖、
并让 seedKeys 在明文下比较。

判据:
- TestRestartDoesNotDuplicateSeededKeys(敏感:回退顺序必红)
- TestSealingIsIdempotentAcrossStarts(12/12 稳定,原先 1/6 flake)
- TestProvidersGetPlaintextCredentials(钉 provider 必须拿到明文这一不变量)
2026-09-28 22:52:51 +08:00
5c58244781 chore(version): 1.7.1 -> 1.7.2
token 统计单位修复(流式改用上游真实 usage、图片不再记 token),
影响 per-model 配额计费口径,走 patch 发布。
v1.7.2
2026-09-28 22:19:24 +08:00
0121d23f91 fix(tokens): 流式统计改用上游真实 usage,图片不再记 token
两处 token 单位错误,均影响 per-model 配额计费:

1. 流式路径的 prompt/completion 只是「字节÷3」估算。
   pumpStream 明明收到了上游最后一帧的真实 usage,却只发给客户端、
   从不回写审计记录,于是配额按估算值扣。生产实测同一请求:
   上游 prompt=37/completion=179 → 记账 27/262,prompt 低估 1.4x、
   completion 高估 1.5x(双向失真)。同模型流式 completion 中位数
   是非流式的 4-27 倍。非流式路径本就用真实值,两路不一致。
   修法:lastUsage 非零时写回 rec.Prompt/rec.Compl,估算降为兜底
   (上游不报 usage 时仍保留原估算行为)。

2. 图片请求把「图片张数」记成 completion_tokens。
   rec.Compl = int64(len(resp.ImageData)),len 是切片长度即张数
   (生产 38 条 image 记录全是 1),且被计入 token 总量。
   图片生成无 token 概念 ⇒ 新增 Req.ImageCount 独立字段,
   Prompt/Compl 归 0;UI 记录表 image 行改显示张数(新增 i18n thImgs)。

顺带补 TestUILocaleKeyParity:此前无人校验 zh/en 键集合一致,
单边加键不会报错,只会显示原始键名。

新增 token_units_test.go(定值上游 6 项),做过变异验证:
回退修复实测复现 stream=16/173 vs chat=44/100、image completion=3。
2026-09-28 22:19:12 +08:00
de7c372ad2 chore(version): 1.7.0 -> 1.7.1
v1.7.0 的配额语义(整钥总额)与最终设计不符,本 patch 版把配额改为
按模型独立计费。已在生产部署过的 v1.7.0 保留不动,语义修正走 patch。
v1.7.1
2026-09-27 19:07:48 +08:00
c51066f0b6 refactor(quota): 配额改为按模型,删除整钥总配额
用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
2026-09-27 19:02:13 +08:00
5530912d32 chore(version): 1.6.0 -> 1.7.0
中版本跃迁:新开 release/v1.7.x 承载 1.7.x 全部 patch。
v1.5.x 已发到 v1.6.0(tag),不再追加。
v1.7.0
2026-09-27 18:46:59 +08:00
cc5e725226 docs: 补齐 per-key 配额的运维视角文档
代码回流 main 时配额小节已随行,但本轮新增的三项认知此前只存在于
commit message 与判据注释里,运维查不到:

1. **配额拒绝 vs 容量拒绝是两种东西**。配额在入口检查、不占上游槽位,
   是廉价拒绝(实测 19–21ms,429 + Retry-After);容量不足要等满
   busyWait 才 503(约 2.6s)。客户端据此可以区分「等窗口重置」与
   「等上游腾容量」——前者只需耐心,后者通常该降并发或换源。
   附实测对照表(容量 8、0.6s/请求、100 并发)。

2. **配额的开销**。每请求检查 149ns(配了配额)/ 42.6ns(未配配额,
   不碰桶)/ 37ns(admin);窗口查询按窗口长度扫描而非扫全量保留
   (1h 49ns、24h 55ns、30d 3.9us);生产形态内存 3.7MB。
   明确写出「未配配额的 key 几乎不付代价」,运维可放心多建 key。

3. **按源 pin 的桶是惰性创建的,以及它的代价**。无条件维护会让
   20 密钥 × 8 模型 × 3 源多占 18MB,所以只有真被 pin 查询过才维护;
   代价是配置 pinned 配额之前的历史用量无法事后按源拆分,首个窗口
   可能少算 —— 这是个会让排障困惑的行为,必须写出来。

同时更新 WebUI 密钥页说明:卡片配额徽标、「配额」编辑按钮、创建表单
可配预算(admin 自动禁用)、「我的密钥」页展示本 key 预算。

README.md / README_EN.md 同步。
2026-09-27 18:46:39 +08:00
652842783f test(gateway): 补配额桶的真实并发竞态判据
per-key 配额桶是共享 map:每个被记录的请求写它,每个配额检查读它。
单线程单测完全看不到这里的竞态,只有让多个 goroutine 同时读写才有效。

key_quota_concurrency_test.go:64 goroutine 跑 2 秒,并发 Record +
KeyWindowTokens + KeyWindowReqs + KeyWindowModelTokens,其中一条路径
在中途 opt in 惰性创建的 pinned 桶(那条路径一次改两个桶 map)。

go test -race 结果:零 DATA RACE,5444 万 token 全部入账。
全仓 -race(./...)亦全绿。

(cherry picked from commit 18cfd6d32b)
2026-09-27 18:44:41 +08:00
a21ae84cbe perf(gateway): 拒绝路径只判定一次 + 补配额交互判据
复查后修掉一个自己引入的缺陷,并补上此前缺失的交叉场景验证。

## 修复:拒绝路径重复判定

4 个入口原本先 checkModelScope(判是否为空)再 writeScopeReject
(内部又 checkQuota 一次)。即每个【被拒】的请求要跑两遍配额统计,
且两次之间用量可能变化 —— 判定与响应存在理论竞态。

改为 checkQuota 一次判定直接把 *quotaRejection 交给 writeReject,
消息与 Retry-After 都来自同一次读,不再有二次求值。
checkModelScope 保留(只需知道放行与否的调用方仍可用)。

## 补判据:此前完全没验证过的交叉场景

1. TestKeyQuotaWinsOverSlotQuota —— key 配额与 AUTO 槽位配额是两种
   不同作用域的限额(槽位是全网关共享的上游预算,key 配额属于单个
   调用方)。两者同时耗尽时必须报【key 配额】:报槽位配额会被表述成
   「无可用容量」,读起来像上游故障,而调用方能处理的恰恰是 key 配额。
2. TestUncappedKeyNeverBlockedByEmptyScope —— 只配模型范围、不配配额的
   key(生产上 5 把 user key 全是这种)绝不能被槽位检查误伤。

## 复查补测的实测数据

配额检查的真实开销(每请求一次,走完整 checkQuota 路径):
  配了配额    149 ns  0 allocs
  未配配额     42.6 ns 0 allocs   <- 生产上 5/7 把 key 是这种
  admin key    37 ns  0 allocs

未配配额的 key 只付 FindKey 的开销、根本不碰桶。相对一次 LLM 请求
(秒级)可忽略。

生产配置副本(7 key / 16 源 / 真加密凭据 / 真上游)实测:
- 100 并发 -> 50 成功 / 50 容量拒绝,RSS 19.9 -> 25.8 MB
- 生产形态桶内存(7 key x 8 model x 2 源 x 40 天满 retention)
  = 3.73 MB,占 ~32MB 预算的 11%
- 配额记账与 stats 一致:配 63000 配额后报 64062/63000
- **跨重启存活**:重启后从审计日志回放,仍报 64062/63000 并拦截;
  未配配额的 key 仍 200

(cherry picked from commit 9811654b3e)
2026-09-27 18:44:41 +08:00
d072a03c9a perf(gateway): 配额桶扫描改为窗口化 + pinned 桶惰性创建
审查本特性线的性能时发现两个问题,均有实测数据。

## 1. 窗口查询是全扫,代价落在每个请求上

sumBuckets 原来遍历整个 map(最多 960 个小时桶),实测 5.9us/op。
配额检查在每个请求上跑 2-3 次(key 总 token、key 请求数、scope token),
于是单请求多付约 18us。

注意这**不是本改动引入的成本**:main 上既有的 WindowTokens 同样是
5907ns/op(全扫)。是本改动让它在请求路径上被调用得更多。

改为只遍历窗口可能覆盖的桶(键是整点小时,范围是精确的,不是采样):
- 24h 窗口 200ns -> 55ns
- 1h  窗口  80ns -> 49ns
- 30d 窗口 5.9us -> 3.9us(720 次查找,只有配 month 配额时才走到)

等价性由 TestSumBucketsMatchesFullScan 保证(400 组随机桶位置 x 6 种
窗口,对全扫逐项比对)。★ 第一次写错成 floor,判据立刻抓到:
30 天窗口报 8878 而全扫是 8649 —— 正确是 ceil。

## 2. pinned 桶无条件创建,内存最坏 26.7MB

每条记录写两个桶:裸 model 与 "source::model"。但 pinned 桶只有
「配额里显式 pin 了 source」时才会被查。

实测最坏情况(20 key x 8 model x 3 source x 40 天全 retention):
HeapAlloc 26.67MB —— 而 README 宣传「16 源生产实例 ~32-35MB」,
等于吃掉 80% 内存预算。

改为惰性:只有 KeyWindowModelTokens 带 source 查询过某个 (key, model)
之后,才开始维护它的 pinned 桶。

  20key x 8model x 3src   26.67MB -> 8.87MB  (-67%)
  5key x 6model(真实)     3.36MB -> 2.26MB  (-33%)
  5key x 12model            5.82MB -> 3.59MB  (-38%)

代价:配 pinned 配额之前发生的用量无法事后按源拆分(记录里虽然有
Source,但桶只存了裸 model),所以 pinned 配额的首个窗口可能少算。
已在代码注释与判据中写明。

## 其余实测

  Record        main 基线 275ns/429B/3allocs -> 283ns/429B/3allocs
                (+8ns,分配数不变;3 allocs 来自 ring buffer)
  配额检查全路径  115ns / 0 allocs(每请求新增)
  纯读路径      6.5ns / 0 allocs

## 100 并发调度/拒绝压测(真实进程 + 可报并发峰值的假上游)

  容量 100(4+96),0.15s/请求,100 并发  ok=100 fail=0   上游峰值 42
  容量 100,0.15s/请求,200 并发          ok=200 fail=0   上游峰值 97
  容量 8,3s/请求,100 并发               ok=8   fail=92  上游峰值 8
  容量 8,0.6s/请求,100 并发             ok=32  fail=68  上游峰值 8
  容量 8,0.6s/请求,40 并发              ok=32  fail=8   上游峰值 8

上游峰值恒定不超过 max_concurrent,容量拒绝返回 503 + busyWait 有界
等待(约 2.6s)。main 基线在同条件下 ok=32 fail=68、上游峰值 8、
延迟分布相同 —— 配额改动没有触碰调度/拒绝路径。

配额拒绝单独验证(低并发避开容量拒绝):req_quota=50 用尽后
100 并发全部 429 rate_limit_exceeded + Retry-After: 1661,
**延迟仅 19-21ms**、上游 total 未增加 —— 配额在入口廉价拒绝,
不占用任何上游槽位,与容量不足的昂贵等待形成明确分工。

## 判据

key_quota_perf_test.go:4 个基准 + 2 个判据(sumBuckets 等价性、
pinned 桶惰性)。3/3 变异全被抓(无条件建 pinned 桶、firstHour 用
floor、keyHour 不再写)。

(cherry picked from commit be11a06a46)
2026-09-27 18:44:41 +08:00
f814ff7468 fix(webui): 修 7 处弹窗关闭错对象 + 模板管理器变量遮蔽
上一提交只修了自己新加的两处弹窗,全站其余 7 处是同一缺陷:所有对话框
共用 id="modal-wrap"(CSS `#modal-wrap:not(:empty){display:flex}`)且可以叠
加(seed-key 提示就盖在密钥页上),而
`const w = $("#modal-wrap"); w.remove()` 移除的是**文档里第一个**,不是用户
刚提交的那一个。

逐处改为两种安全写法:
- 能拿到按钮的(saveSource / saveTemplate / sortScopeSave / scopeSave /
  keyQuotaSave / createKey / downloadStatsCsv / downloadKeysCsv /
  scrAddFromForm):`btn.closest("#modal-wrap")`
- 拿不到按钮的:新增 `closeTopModal()` 取**最后一个**(用户看到的那个),
  并作为所有 `if (w) w.remove()` 之后的兜底
- 顺带把 sortScopeSave / downloadStatsCsv / downloadKeysCsv / scrAddFromForm
  的签名补上 btn / this 参数 —— 否则 .closest 恒为 null,表单永远不关

同时修一个相邻的既有 bug:`openTemplateModal` 的 `.map((t) => ...)` 用 t 做
循环变量,模板字面量里又调 t("srcEdit"),t 被遮蔽成对象 ⇒ 打开模板管理器
直接抛 `t is not a function`,整个弹窗渲染失败(main 上就有,git show 确认)。
参数改名 tpl。修后模板管理器完整渲染(浏览器实测:DeepSeek / 智谱 / Kimi /
SiliconFlow 各行 + Edit/Delete 按钮文案全部正常,零异常)。

判据从 2 处扩到全量(internal/gateway/ui_quota_contract_test.go,+3 例):
- 全文档扫描:任何 `.remove()` 配裸 `$("#modal-wrap")` 即失败
- 9 个关闭对话框的处理器必须走 .closest 或 closeTopModal
- closeTopModal 必须取 all.length - 1(首尾颠倒就是原 bug)
- 用 .closest 的处理器,其签名必须真的有 btn 参数 —— 否则查找恒为 null,
  表单永远不关
5/5 变异全被抓:sortScopeSave 去 btn 参数、closeTopModal 取第一个、scopeSave
退回裸选择器、saveSource 退回裸选择器、downloadKeysCsv 删兜底。

浏览器实测(共享 Chromium CDP,真实进程,每处都插入一个「decoy」弹窗
占据文档首位,复现原 bug 的触发条件):
- scopeSave / keyQuotaSave / scrAddFromForm / saveSource / saveTemplate
  五个处理器:自己的表单关、decoy 保留 ✓
- 零 JS 异常
★ 测试自身踩了两个坑,都不是代码问题:① `document.querySelector(sel) && .click()`
  在 CDP 里求值为 undefined,改成箭头函数;② 编辑模板时没填名字就点保存,
  saveTemplate 因 `if (!nm)` 早退、fetch 零调用 —— 一开始我把这个误读成
  「修复失效」,加 fetch 拦截 + 读 #s-name 的值才定位到是测试数据缺失。
  ⇒ 「点按钮没反应」要先分清是「事件没触发」「请求失败」还是「早退」。

(cherry picked from commit b5c3fea0bb)
2026-09-27 18:44:41 +08:00
ef631b43dd feat(webui): 密钥配额表单 + 修复弹窗关闭错对象
功能:让 per-key 配额在 WebUI 里可配置可见,之前的实现只有 API 与
config.yaml 能配。

- 密钥卡片头部显示配额徽标(token / 请求数 + 重置窗口),admin key
  不显示编辑入口(服务端本就永不受限,给入口只会让人以为配了会生效)。
- 新增「配额」编辑弹窗:token 配额、请求数配额、重置周期(复用既有
  的 period 词表与 n-hour 联动),预填从 canvas 的 data-* 读。
- 创建密钥弹窗同步加配额字段;选 admin 角色时自动禁用(同样因为服务端
  忽略 admin 的配额)。
- 「我的密钥」页新增 KEY-WIDE QUOTA 列,用户能看到自己这把 key 的预算。

修一个真 bug:保存弹窗用 $("#modal-wrap") 关闭自己,而全站弹窗共用这个
id、且可以叠加(seed key 提示就盖在密钥页上)。实测(共享 Chromium
CDP,seed 提示与配额弹窗共存)确认:保存后被移除的是 seed 提示,配额表单
反而留在屏幕上 —— 症状是「保存了但弹窗没关」,指向的方向完全错。改为用
点击的按钮 btn.closest("#modal-wrap") 解析自己的弹窗。createKey 有同样
问题,一并修。既有文件里另有 7 处同样写法,未动(不在本次范围,且新判据
只对本次改的两处断言,避免误伤)。

判据新增 internal/gateway/ui_quota_contract_test.go(6 例):
- 两个表单必须用 .closest 解析自己的弹窗
- putScope 必须带上 4 个配额字段(API 视其为指针,省略=清空预算)
- 创建请求必须真的发出配额字段
- **数据流判据**:徽标要真读 k.token_quota 等、编辑表单要真读
  canvas 写的 data-kquota 等。只查字面量存在会漏 —— 字段躺在死分支里
  判据照样通过(这是本轮实际踩到的:keyCapBadges 经 keyPeriodSuffix
  间接读 k.period,被判据抓到后我把读取显式化而不是放宽判据)
- 弹窗扫描先剥注释,否则修复说明里引用的字面量会被当成违规
- 复用既有 ui_contract_test.go 的 jsFunctionBody(大括号配平);
  自己第一版用 2000 字符固定窗口,被长注释顶开后仍在窗口外命中后面
  函数的同名字段,读起来像通过 —— 窗口法在这里是假判据

7 个变异全部被抓(unsafe 关闭、putScope 丢字段、createKey 丢字段、
徽标不读字段、canvas 不写 data-*、kq-hours 改名、周期词表缺项)。

浏览器实测(共享 Chromium CDP,真实进程 + 加密配置):
- 徽标渲染 1.0K·1h / 5×·1h;编辑框预填 1000/5/hour,hours 框按周期联动
- 保存后回读 250000/77/nhour/6,徽标更新为 250.0K·6h,toast Saved
- 零 JS 异常
- **关键回归**:编辑模型 scope 后配额仍是 250000/77/nhour,未被清空
- 创建带配额的 key,服务端确认 {t:50000,r:300,p:week,role:user}
- user 视角「我的密钥」显示 777·1h 与 9×·1h

文档:README.md / README_EN.md 补「密钥用量配额」小节(配置示例、
周期词表、429 语义、admin 豁免、整点分桶最晚晚 1 小时释放、PUT 的
省略 vs 0 语义、429 响应样例),特性列表各加一条。

(cherry picked from commit ce66c7f6c2)
2026-09-27 18:44:41 +08:00
9c3aabb7f9 feat(gateway): per-key 用量配额(token + 请求数)与重置周期
问题:密钥控制只能限制模型范围。实测发现三个缺陷,其中前两个让
per-model token_quota 在真实链路上从未生效:

1. 桶键不含 key。scopeTokens 调 WindowTokens(model, source, win),
   桶键是 model / source::model,与调用方无关。实测两把 key 各用
   1000 token,窗口报 2000 —— A key 的额度被 B key 消耗。
2. 无 source pin 的桶永远是空的。真实记录 Source 总被填上,桶键存成
   "deepseek::m1",而无 pin 的查询找 "m1" —— 读到 0,永远 < quota,
   配额形同虚设。实测 WindowTokens("m1","",1h)=0 而 pinned=2000。
3. AUTO scope 走 KeyTokens(key),是全时段累计、永不重置。实测 30 天
   前的 200 token 仍计入 1 小时配额(报 210 而非 10)。配了
   period: hour 也不会每小时归零。

生产 5 把 user key 全是 token_quota: 0,所以前两条一直没暴露。

改动:
- Stats 新增 per-key 小时桶 keyModelHour(key → model → hour)与
  keyHour(key 总量)、keyReqHour(请求数),retention 40 天,与既有
  modelHour 对齐以覆盖最长的 month 窗口;LoadAudit 走 aggregateLocked,
  所以窗口用量跨重启存活。modelHour 保持 key-blind:它服务的是 AUTO
  槽位配额(限制整个网关对某槽位的消耗),语义不同,不应被 per-key
  改造污染。
- 每个请求写两份模型桶:裸 model 与 source::model。无 pin 的 scope
  条目读前者,有 pin 的读后者。
- GWKey 新增 TokenQuota / ReqQuota / Period / Hours:整钥配额,
  跨该 key 所有模型共享一份预算;ReqQuota 覆盖持续请求量(源上的
  RPM 只管突发)。
- 配额耗尽返回 429 + Retry-After(rate_limit_exceeded),而不是 403:
  403 让客户端以为这把 key 永远不能用该模型,直接放弃;429 + 等待
  才能在窗口重置后自动恢复。模型越权仍是 403。
- admin key 永不受配额限制 —— 否则操作者会把自己锁在门外。
- 周期词表在写入时校验,拼错的 period 被拒绝而不是静默当成永不过期
  (那与操作者输入的意图正好相反)。
- PUT /api/keys 的配额字段是指针:省略=保留原值,显式 0=解除限制。
  否则只改模型范围就会悄悄清空预算。

判据 3 个文件 24 例,9 个变异全部被抓:key 隔离、pin 桶缺失、
AUTO 周期、key-blind 退化、429→403、admin 被限、PUT 清空配额、
Validate 失效、pinned 桶缺失。前三个变异最初漏网 —— 判据只测了
Stats 层没测接线,补了走真实 HTTP 的接线层与 API 层判据后抓住。
端到端验证:真实进程 + 加密配置往返,配额字段与 enc:v1 密钥均正常。

(cherry picked from commit 5306251840)
2026-09-27 18:44:41 +08:00
a7355debed feat(deploy): 部署前校验 master.key 可解封配置 + llmsproxy -show-secrets
密钥校验(deploy.sh)
- 新增 verify_master_key,在 build/替换任何文件之前执行。失败则二进制与
  配置分毫未动、服务不受影响(已负向验证:缺钥匙、错钥匙两种情况都挡住)
- 走真实的 -show-secrets 解密路径,而不是只检查钥匙文件格式——格式合法
  但内容不匹配(重新生成、恢复了错的备份、换机器)同样会被拒
- 钥匙来源与 config 包一致:LLMS_PROXY_MASTER_KEY 优先,否则
  dirname(runtime_file)/master.key
- 配置里没有密文时跳过并提示(首次加密场景)

-show-secrets
- llmsproxy -show-secrets -config <path>:把凭据打到 stdout 后退出
- 不启动任何东西、不写任何文件(已验证 mtime 不变)
- 加密往返无损:封存前后输出逐字节一致

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit a8cff57e24)
2026-09-27 17:12:03 +08:00
f6baa13583 feat: WebUI 改为依赖 /api/v1,UI 与 agent 共用一套 API 契约
- sources / sort / keys 三个页面的数据源从 /api/sources 切到 /api/v1/sources
  (写操作仍走 /api/sources:v1 是只读门面,不做变更)
- 编辑弹窗改用 /api/v1/sources/{name}?reveal=credentials(admin-only)取明文 key。
  这是必须的:表单要整体回传源,若不回填 key,改个端口就会把 key 清空。
- 遮蔽视图仍是默认,只有显式 reveal 才返回明文

端到端验证(真浏览器 + 临时实例,非仅 API 测试):
- sources/sort/keys 三页实际发出 GET /api/v1/sources,0 console error
- editSource('demo') → reveal=credentials,#s-key 与 #s-url 正确回填
- 写入往返:改 base_url /v1→/v2 后重开,key 仍在(未被清空)
- 落盘 api_key 明文残留 0、密文 1

测试:+1(reveal 必须 admin,否则任意 user key 可读全部凭据)
变异验证:reveal 去掉 admin 校验 → 403 断言变红

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 926b9f6565)
2026-09-27 17:12:03 +08:00
fd03ef6e4a feat: 密钥静态加密 + /api/v1 agent 管理 API
密钥加密(写侧封存 / 读侧解封)
- config.yaml 的 sources[].api_key、sources[].headers、keys[].key 落盘即
  AES-256-GCM 密文(enc:v1: 前缀),master.key 复用 runtime store 那把
- 内存里永远是明文:鉴权比对、API 返回新建 key、WebUI 编辑回填都不受影响
- 启动时一次性封存现存明文(幂等,已封存则不写盘);-check 不写文件
- UpsertSourceInYAML 增加 box 参数,新加的源不再以明文落盘
- 解密失败改为硬错误:原先 MustDecrypt 返回密文会被下次 Save 二次封存
  (实测:源 key 18→20、静默损坏),现在启动即失败且配置分毫不动

/api/v1:面向 agent 的管理 API(WebUI 零影响)
- GET /api/v1            机器可读索引,列出每个端点的方法/权限/用途
- GET /api/v1/overview   一次调用看全貌:源 + AUTO 链 + 密钥数 + 健康度
- GET /api/v1/health     仅健康快照
- GET /api/v1/models     按源分组的可路由模型清单
- GET /api/v1/sources[/{name}]  凭据遮蔽后的源
- GET /api/v1/auto       调度链与实时槽位状态
- GET /api/v1/keys       admin only,密钥元数据,绝不回显密钥本身
- 沿用同一套网关 key 鉴权;读端点任意角色,写仍需 admin

测试:15 个新用例(含负向:泄密、越权、写操作必须被拒)
变异验证:maskKey 不遮蔽→红、去掉 admin 校验→红

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit ad28a924a5)
2026-09-27 17:12:03 +08:00
97bb9c6ef4 feat(opencode): 透传 completion_tokens_details.reasoning_tokens 与上游 cost
回答「opencodego 的用量与费用透传呢」时逐字段核对上游产出,发现 usage 漏了
一项、费用整项丢失。

## 上游实际发什么(实测 opencode.ai/zen/go/v1)

  {
    "choices": [...],
    "usage": { "prompt_tokens": 37, "completion_tokens": 40, "total_tokens": 77,
               "prompt_cache_hit_tokens": 0, "prompt_cache_miss_tokens": 37,
               "prompt_tokens_details": {"cached_tokens": 0},
               "completion_tokens_details": {"reasoning_tokens": 40} },
    "cost": "0"
  }

cost 在**顶层**且是**字符串**。流式时还会单独发一帧:
{"choices":[],"cost":"0"}

## 此前丢了两样

1. completion_tokens_details.reasoning_tokens —— 输出里有多少是思考 token。
   没有它,客户端无法判断 completion_tokens 里多少是可见回答、多少是思考,
   而两者都按输出计费。
2. cost —— 唯一的费用信号,网关整个丢弃。Go 订阅是包月制恒为 "0",
   但 Zen 按量付费模型(以及未来的其它源)有信息量。

顺带修掉一处流式/非流式不一致:命中缓存时上游同时给
prompt_tokens_details.cached_tokens 和独立的 hit/miss,流式路径写成了 elseif,
只留 details,与非流式产出不同(只认独立字段的老客户端会看不到缓存)。

## 实现

- types.TokenUsage += CompletionTokensDetails;UnifiedResponse / UnifiedChunk += Cost
- opencodego/opencodezen 适配器映射两个字段;空 choices 帧改成 usage 与 cost
  都可带(早退只带 usage 会把同帧的 cost 丢干净 —— 新测试先抓到的就是这个)
- Gateway ChatCompletion / ChatChunk += cost,随终帧发(对齐上游的
  {"choices":[],"cost":"0"} 形态)
- Go 兜底 standardSSEChunk 同步支持(openai 系适配器不再漏 reasoning_tokens;
  纯 cost 帧不再被整体丢弃),新增 rawCostString 兼容字符串/数字两种形态

费用只做**搬运**:不解析、不换算、不汇总 —— 它是上游事实,且只有部分上游提供。

## 验证

经网关实测 gozen:deepseek-v4.1-flash,流式与非流式产出逐字段一致:
  prompt_tokens_details.cached_tokens=6784
  prompt_cache_hit_tokens=6784 / miss=148
  completion_tokens_details.reasoning_tokens=16
  cost="0"

测试:TestOpenCodeCostAndReasoningPassthrough(含「无数据不得凭空造字段」反例)、
TestOpenCodeStreamCacheFieldsMatchNonStream、TestTokenUsageMarshalsCompletionTokensDetails。

(cherry picked from commit c744ee151e)
2026-09-27 17:12:03 +08:00
d81074621a fix(opencode): 采纳客户端真实会话 id + 超窗消息不再被限流措辞封杀
两处都源于同一次排查:pi 到底有没有带会话标识、超窗为什么触发不了压缩。

## 1) 客户端会话 id:pi 一直在发,只是被配置关掉了

之前结论是「通用客户端不发会话 id」——只对了一半。pi 有会话 id,且能发:
pi-ai 的 createClient 在 compat.sendSessionAffinityHeaders 为真时,会把
平台会话 id(uuidv7,整个会话恒定)放到 x-session-affinity /
x-client-request-id / session_id 上。该开关默认 false,而 llmsproxy 的
provider 配置里没开,所以此前一直收不到。

现在网关按优先级采纳:x-session-affinity → x-session-id → session_id →
body 的 prompt_cache_key,并把值经 types.ChatRequest.ClientSession 传到
适配器 meta.client_session。适配器的会号种子优先级变为:
客户端会话 id > 首条 user 消息指纹 > 按源固定。

刻意不采纳 x-client-request-id:名字含 request,部分客户端每请求都换,
拿它当会话会让上游前缀缓存永不命中(pi 总会同时发 x-session-affinity,够用)。

实测:抓 127.0.0.1:8081 的真实 pi 请求,配置打开后收到
x-session-affinity = session_id = x-client-request-id = <子会话 uuid>。
上游缓存确为会话级隔离(同前缀、不同会号:A 冷→命中,B 首次仍为 0),
两个不同 header 值互不命中,反证网关确实采纳了客户端会话 id。

## 2) 超窗消息必须「干净」,否则被同链的限流措辞反向封杀

pi 的 isContextOverflow 先查 NON_OVERFLOW_PATTERNS(/rate limit/、
/too many requests/、Bedrock 前缀),命中就直接判为「非超窗」——**即使
消息里已经有 context_length_exceeded**,pi 也不会压缩重试。

而 AUTO 链的失败消息天生是多 tier 原因的拼接,超窗 tier(gozen 400
maximum context length)常与配额/限流 tier(429 token plan exhausted、
cooling、no free slot)同时出现。此前把 tier 明细原样拼在归一化标记后面,
等于让一条限流 tier 的措辞反过来封杀超窗识别。

现在超窗走独立的干净消息:
  context_length_exceeded: context window is full; reduce the length of
  the messages (gozen/deepseek-v4.1-flash)
只留超窗措辞 + 超窗源名,不带任何其它 tier 的文本。

测试:TestOverflowMessageSurvivesRateLimitedSiblingTier 用 pi 的完整判定
顺序(先 NON_OVERFLOW 后 OVERFLOW)断言同链限流 tier 不再封杀超窗识别;
TestClientSessionFromRequestHeaders / TestClientRequestIDIsNotUsedAsSession /
TestOpenCodePrefersClientSessionID 覆盖会话采纳与优先级。

(cherry picked from commit c9c09b2ba2)
2026-09-27 17:12:03 +08:00
4dd3431c26 feat(opencode): per-conversation session via first-user-message fingerprint
Follow-on to the session-stability fix. "Per source" already made the
prefix cache hit, but it puts every conversation into one upstream session.

Using the client's own session id is not possible: capturing real agent
traffic (tcpdump on 127.0.0.1:8081) shows generic OpenAI clients send NO
session identifier at all — no user / session_id / conversation_id /
metadata in the body, and no session header (only X-Stainless-* plus
User-Agent: pi). The x-opencode-session the Go endpoint asks for is an
OpenCode native-client concept that a generic client cannot forward.

Since history is replayed every turn, the FIRST user message is invariant
for the life of a conversation, so it is used as the conversation
fingerprint. The session becomes stable within a conversation and distinct
across conversations; requests with no user message fall back to per-source
stability.

Measured through the gateway (same 5.7k-token prompt): 2nd call
cached_tokens=5504, and an unrelated conversation gets its own session.

Test: TestOpenCodeSessionIsStableForCache covers same-conversation
stability, cross-conversation separation, per-request request ids and the
sessionless fallback.

(cherry picked from commit 39b48e556e)
2026-09-27 17:12:03 +08:00
3d1407e6fa fix(opencode): make x-opencode-session stable so the upstream prefix cache can hit
The opencode adapters derived x-opencode-session from meta.timestamp, i.e. a
brand new session on every request. The upstream prefix cache is
session-scoped, so no request could ever hit it, and the cache fields the
endpoint does report (prompt_tokens_details.cached_tokens,
prompt_cache_hit_tokens/prompt_cache_miss_tokens) always came back 0/absent.

Measured against the live endpoint, same 6032-token prompt:

  fixed session id   -> 2nd call: hit 5888, miss 144
  rotating session id -> every call: hit 0, miss 6032

Fix: derive the session from the source name (stable), matching how
x-opencode-project is already derived. x-opencode-request stays unique per
request — it is only a request identifier, not part of the cache key.
Applied to both opencodego and opencodezen.

Through the gateway the same prompt now reports, on the 2nd call:
  details={'cached_tokens': 5888} hit=5888 miss=144      (non-streaming)
  prompt_tokens_details={'cached_tokens': 5888}          (streaming)

Test: TestOpenCodeSessionIsStableForCache asserts the session is stable
across requests for one source while the request id differs.

(cherry picked from commit 791d198f47)
2026-09-27 17:12:03 +08:00
db320d2f6c fix(opencodego): inject empty reasoning_content on tool-calling turns
v1.5.4 stopped stripping reasoning_content, which fixes clients that send
it — but most agent clients (pi included) never store or replay their
reasoning, keeping only the tool call. OpenCode Go validates the field on
any assistant turn that carries tool_calls and rejects the whole request:

  400 invalid_request_error: The `reasoning_content` in the thinking mode
  must be passed back to the API.

Verified against the live endpoint that an EMPTY string satisfies the
check, so the adapter now fills in "" when a tool-calling assistant turn
has no reasoning_content. Nothing is fabricated: the reasoning shown to
the client is still exactly what the upstream returned for that turn.

Measured: with a tool_call + tool_result history and no reasoning_content,
all 25 configured Go models returned 400 before and all 25 answer
correctly now.

Test: TestOpenCodeGoVsZenReasoning also pins that a plain assistant turn
(no tool calls) must NOT gain the field.

(cherry picked from commit d1a72cd23a)
2026-09-27 17:12:03 +08:00
b544671f5e feat(adapters): split opencode into opencodezen and opencodego
Zen (https://opencode.ai/zen/v1) and Go (https://opencode.ai/zen/go/v1)
are different services with different requirements, and one shared adapter
could not satisfy both.

The decisive difference is reasoning_content:

  * OpenCode Go runs thinking models and REQUIRES the assistant turn's
    reasoning_content to be echoed back. The shared adapter stripped it
    (msg.reasoning_content = nil), so every replay of a thinking turn
    failed with:
      400 invalid_request_error: The `reasoning_content` in the thinking
      mode must be passed back to the API.
    Reproduced directly: the same request with reasoning_content -> 200,
    without -> 400. That is why the Go tier never worked in an agent loop.

  * The Zen free pool must not receive it, so it keeps stripping.

Both adapters keep the earlier fixes they share (never drop an assistant
turn carrying tool_calls; send stream_options only when streaming; role
whitelist; multimodal strip) and the opencode client fingerprint headers —
the Go endpoint additionally REQUIRES x-opencode-session, which the
adapter already sends.

config: localzen -> opencodezen, gozen -> opencodego.
Verified: all 25 gozen models answer correctly through the gateway with a
thinking + tool_call + tool_result history (was 0/25 before), streaming
included; the Zen free models still pass.

Test: TestOpenCodeGoVsZenReasoning pins the Go-keeps / Zen-strips split.
(cherry picked from commit 71e9040a45)
2026-09-27 17:12:03 +08:00
0881a49286 fix(opencode): only send stream_options with stream:true
OpenCode Go (and other strict OpenAI-compatible upstreams) reject a
non-streaming request that carries stream_options with
"stream_options should be set along with stream". The adapter attached it
unconditionally, so every non-stream call through the opencode adapter
failed on those upstreams.

Verified against OpenCode Go: 25/25 configured models now pass a real
completion through the gateway (they previously 400'd).

Test: TestOpenCodeStreamOptionsOnlyWhenStreaming (absent when
non-streaming, include_usage present when streaming).

(cherry picked from commit 0528941e24)
2026-09-27 17:12:03 +08:00
4e2b5bfda9 fix: empty-array content becomes invalid {} on every pass-through adapter; gemini/ollama drop tool calls
Three related forwarding defects found by auditing every adapter with a
tool-calling replay (assistant turn with content:[] + tool_calls).

1) content:[] -> content:{} (all 12 openai-adapter sources, plus
   deepseek/trae/sensenova/agentrouter/github/groq/kimicode/mistral)

   Lua adapters json.decode the request and re-encode it, and an empty Lua
   table is indistinguishable from an empty JSON array — the encoder emits
   {} for both. Agent clients serialise a tool-calling assistant turn with
   no text as content:[], so every pass-through adapter rewrote it to
   content:{} — not valid OpenAI (content is string|array|null). Verified
   against a live upstream: content:[] produced "400 invalid arguments"
   while content:"" was accepted.

   Fixed once at the decode boundary (types.ChatMessage.UnmarshalJSON):
   empty-array content normalises to "" and an empty tool_calls array is
   dropped, so every adapter — including future ones — sees a valid shape.

2) gemini dropped tool_calls and never emitted functionCall /
   functionResponse; the tool role also stayed as an invalid role inside
   contents and system was not moved to systemInstruction.

3) ollama copied only role/content, dropping tool_calls and the call
   attribution entirely (it needs tool_name, not tool_call_id).

Test: TestAdaptersPreserveToolCalls asserts, for every adapter, that the
call id (or function name where the wire format has no id), the function
name, the tool result and the trailing user turn all survive, plus a
negative control for plain text.

(cherry picked from commit 9114468753)
2026-09-27 17:12:03 +08:00
61b847da86 fix(opencode): never drop an assistant turn that carries tool_calls
An agent client (pi) serialises an assistant turn whose content is only
[thinking, toolCall] as content:[] with tool_calls. The multimodal-strip
pass treated an empty content array as 'nothing left, drop the whole
message' and discarded the tool_calls with it.

The next message is that call's tool result, so it arrived orphaned: the
model saw a result for a call it had never made and re-issued the same
call on every turn — an endless repeated-tool-call loop. Reproduced
against a capture sink: content:[] lost tool_calls, while content:"" and
content:null kept them.

Only messages with neither usable content NOR a tool call now get
dropped. Content that collapses to empty but still has tool_calls or a
tool_call_id is emitted as "" instead.

Test: TestOpenCodeKeepsToolCallWithEmptyContent (plus a negative control
that an image-only message without tool calls is still dropped).

(cherry picked from commit 499f0cac2f)
2026-09-27 17:12:03 +08:00
6432b61eb6 fix(gateway): AUTO scope grants all models — restrict to routing mode only
A key with scope=[AUTO] could previously:
1. request ANY concrete model id directly (hasScopeModel/checkModelScope
   treated AUTO as a wildcard)
2. see the full 56-model list on /v1/models (intersectModels considered
   AUTO as grant-everything)

AUTO now only authorizes the AUTO routing mode. Direct requests to a
specific model require an explicit scope entry.

Also carries agentrouter.lua WAF fingerprint headers (Origin/Referer/
X-Requested-With) already staged on this branch.

Tests: TestHasScopeModelWithSourcePrefix updated; full suite green.
(cherry picked from commit 7fb8f96b82)
2026-09-27 17:12:03 +08:00
dd708c9680 fix(gateway): 上下文超窗错误归一化,客户端才能压缩重试
问题:上游返回上下文超窗时,客户端(pi)既不压缩也不重试,只看到一条
普通上游错误。链路两处叠加:

1. 措辞不在客户端识别列表里。pi 靠 @earendil-works/pi-ai 的
   OVERFLOW_PATTERNS(25 条正则)判断超窗,而 justworker 返回的是
   「请精简对话历史…(Context window is full…)」——与那 25 条一条都不匹配
   (最接近的 /context window exceeds limit/i 也不命中,因为措辞是 "is full"
   而非 "exceeds limit")。
2. AUTO 链路的每条 tier 错误按 80 字节 OneLine 截断,而 "Context window
   is full" 这类短诊断词常出现在尾部,正好被按字节切掉(直连路径是 160
   才保住)。

修法:
- 新增 overflowMarkers 识别超窗措辞(含中文写法),命中时把客户端可见
  错误归一化为 context_length_exceeded 前缀——它命中 pi 的
  /context[_ ]length[_ ]exceeded/i,超窗因此可被发现并触发压缩重试。
- AUTO 链路每条 tier 错误宽度 80→160,短诊断词不再被截断。

归一化只加前缀,原始诊断信息保留,便于定位是哪一层超窗。

测试:internal/gateway/overflow_err_test.go(5 例,含「标记必须命中 pi
正则」、非超窗不得误标、80 vs 160 宽度的回归对比)。
2026-09-10 20:38:39 +08:00
f0986751de Merge feature/agentrouter-id-sanitize: 补上最后一个服务 Claude 的适配器
全仓审计发现 agentrouter 也暴露 claude-opus-4-8,却是唯一没有 id 消毒的
Claude 适配器。补齐后,所有含 claude/opus/sonnet 模型的源(qijiar、toter、
juziai、agentrouter、justwoker、api456、tabitoken)全部走消毒路径;
测试新增三适配器结果一致性断言。
2026-09-06 10:08:52 +08:00
c6c3e0dcd7 fix(agentrouter): sanitize tool-call ids — it fronts Claude too
Full-repo audit after the anthropic/openai fix: agentrouter exposes
claude-opus-4-8, so it inherits Anthropic's tool id rule
^[a-zA-Z0-9_-]{1,64}$ and rejects the whole request on a violation, exactly
like justwoker/tabitoken/扇贝. It was the only remaining adapter serving Claude
models without the sanitizer, so a client that had picked up a dirty id (e.g.
"bash:0" from moonshotai/kimi-k3) would still lose every turn here.

Same shape as the other two — inbound tool_calls[].id + tool_call_id, outbound
non-streaming ids and the first streamed fragment — and the test now asserts
all three adapters rewrite an identical input identically, so a client mixing
sources within one session cannot end up with unpaired tool calls.

Audit result: every source exposing a claude/opus/sonnet model (qijiar, toter,
juziai, agentrouter, justwoker, api456, tabitoken) now routes through a
sanitizing adapter.
2026-09-06 10:08:52 +08:00
b9944332ff Merge feature/anthropic-usage-cache: 缓存命中不该让 prompt_tokens 算错
Anthropic 的 input_tokens 不含缓存部分,OpenAI 的 prompt_tokens 含。
直接对映射等于漏算全部缓存 token,还能算出 >100% 的命中率;
cache_creation 完全没读;流式路径连缓存字段都整个丢掉。
三项相加 + 共享 map_usage,未上报缓存的上游输出逐字节不变。
2026-09-06 09:55:22 +08:00
6bdb9fcc44 fix(anthropic): count cached input in prompt_tokens instead of dropping it
Anthropic and OpenAI disagree on what the prompt count means:

  Anthropic: input_tokens EXCLUDES cached blocks; cache_read_input_tokens and
             cache_creation_input_tokens are separate, additive, billed input.
  OpenAI:    prompt_tokens INCLUDES its cached_tokens subset.

anthropic.lua mapped input_tokens straight onto prompt, so a cache-heavy turn
was doubly wrong: the billed prompt was undercounted by the entire cache
portion, and cached_tokens could exceed prompt_tokens — a cache hit rate above
100% for any client that divides one by the other. cache_creation_input_tokens
was never read at all, so a cache-write turn silently lost those billed tokens.

Worse, the streaming path dropped the cache split entirely: message_delta
carries the FINAL usage and only mapped input/output, so every streamed
response reported no cache information even when the upstream sent it.

All three counts are now summed into prompt, with the read half exposed as
prompt_tokens_details.cached_tokens plus the DeepSeek-legacy hit/miss pair, via
one shared map_usage() used by transform_response, message_start and
message_delta. A reported zero stays distinguishable from "never reported": the
split is emitted whenever either cache field is present, and omitted entirely
when the upstream mentions neither (justwoker reports only input/output plus its
own cost fields, so its output is byte-identical to before). map_usage returns
nil for a countless object, preserving "no usage in this chunk means say
nothing" rather than reporting zeros.

message_start's placeholder count is still emitted: justwoker reports 160 there
and the real 6931 in message_delta, and the gateway's mergeUsage lets the later
non-zero value win.
2026-09-06 09:55:22 +08:00
25eb30c15b Merge feature/toolcall-id-sanitize: 一个上游的脏 tool_call id 不该拖垮整条 Claude 链
xinjianya/moonshotai/kimi-k3 会吐出 'bash:0' 这种 id。OpenAI 不校验,
Anthropic 校验 ^[a-zA-Z0-9_-]{1,64}$ 且整条请求直接 400。客户端会把它存进
历史并回放给所有源,于是 justwoker / tabitoken / 扇贝(都是 Claude 上游)
同时挂掉,AUTO 一路穿透四个 tier。

两个适配器、进出双向消毒;已合法的 id 逐字节透传;流式参数分片保持无 id。
2026-09-05 22:19:15 +08:00
131a42a169 fix(adapters): sanitize tool-call ids so one bad upstream can't kill every Claude slot
Anthropic requires tool_use.id / tool_result.tool_use_id to match
^[a-zA-Z0-9_-]{1,64}$ and rejects the WHOLE request otherwise with
REQUEST_BODY_INVALID / "Invalid tool use format". OpenAI has no such rule, so
an OpenAI-compatible model can mint an id like "bash:0"
(xinjianya/moonshotai/kimi-k3 does exactly that).

In a fan-out router that id does not stay local: the client stores it in its
history and replays it to every other source. One such id therefore kills
every Claude slot at once — justwoker, tabitoken and 扇贝 are all
Claude-behind-{OpenAI,Anthropic} — and an AUTO request falls through all four
tiers to whatever tolerant model is left. Observed live: 4 consecutive 503
"all N auto providers failed" with tier 1/2/3 each reporting the same 400.

Both directions are sanitized, in both adapters:
  - request:  tool_calls[].id and tool_call_id, so poisoned history recovers
  - response: non-streaming tool_calls[].id and the first streamed fragment,
              so a bad id never enters a client session in the first place

safe_tool_id is pure and deterministic, so a call and its result are rewritten
identically within one request. A rewritten id keeps an 8-hex digest of the
original, without which distinct ids could collapse ("a:b" and "a_b") into a
duplicate/unpaired tool_use. Already-legal ids pass through byte-identical, so
well-behaved traffic is unaffected. openai.lua carries its own copy because
Lua adapters have no shared prelude.

Streamed argument fragments carry no id and must stay id-less, otherwise
index-based accumulation on the client breaks; a test pins that.
2026-09-05 22:18:31 +08:00
3cdb16c906 Merge feature/deploy-config-rollback: 部署脚本把配置纳入回滚闭环
上一轮 446 次重启循环的根因不是代码,是 deploy.sh 把 config.yaml 当作
部署范围之外的东西:健康检查失败后只回滚二进制,坏配置仍在,于是旧
二进制照样解析不了,服务在 systemd 重启循环里空转。

修法是把配置变成一等部署对象:--config 投放 + 重启前 -check 预检 +
回滚时二进制与配置一起恢复。预检位置比预检本身更重要——排在 restart
之前,坏配置才会在服务仍健康时被拦下。
2026-09-05 09:47:04 +08:00
ef92ce4d82 Merge feature/source-proxy: 让被 CF/格式问题挡住的源真正可用
两件事都在回答同一个问题——为什么配了源却用不了:

- trae 源在流式下把模型打印的 tool_call 当纯文本透传,客户端收到
  finish_reason "stop" + 一段文本,agent 循环当场终止
- justwoker/tabitoken 坐在 Cloudflare 后面,直连 403,必须走代理 +
  浏览器 UA;但全局 env 代理会把 trae/localzen 等内网源也绕出去

因此引入 per-source proxy_url:指定的源走专属代理,未指定的保持直连。
2026-09-05 09:46:54 +08:00
c524831616 fix(deploy): roll back config together with the binary, and preflight it before restart
The 446-restart-loop incident: adding a headers block to a source without
removing the source's existing `headers: {}` produced a duplicate YAML key.
The process exited on startup, healthcheck failed, and rollback restored only
the binary — so the old binary kept parsing the same broken config and the
service span in systemd's restart loop. Config was treated as out of scope
for deployment; it is not.

Three changes close the loop:

1. cmd/llmsproxy: new `-check` flag validates a config (parse +
   ApplyDefaults) and exits, without starting the Lua VM, touching
   runtime.json, or binding a port — safe to run against a live service.
   Unlike normal startup it does NOT create a default config, so a missing
   file is an error.

2. deploy.sh `--config <file>`: stage a config for deployment, atomically
   renamed into place with the same copy->rename(2) technique as the binary.
   Omitted means the live config is left alone.

3. Ordering: config replacement and preflight both run BEFORE
   restart_service, so an invalid config is caught while the service is still
   healthy and never triggers a restart. rollback() now restores binary AND
   config (only when this run replaced it, so concurrent WebUI edits survive),
   then re-runs -check before restarting — refusing to restart into a config
   that still fails, instead of trading one restart storm for another.

Also: the sha256-unchanged early exit now only fires when there is no pending
config, otherwise `--config` would be silently dropped.

Verified on the live deployment:
- reproduced the exact duplicate-key config: preflight caught it, PID
  unchanged (zero interruption), binary and config both rolled back, gateway
  still answering 200
- valid config: replaced, service restarted, new value live
- no --config: binary-only deploy unaffected
- go test -tags luajit ./... passes
2026-09-05 09:46:43 +08:00
868ac59692 feat(provider): support per-source proxy_url for geo-blocked upstreams
Upstreams like justwoker/tabitoken sit behind Cloudflare geo/IP blocks and
only respond through a proxy. A global env proxy is wrong (intranet sources
trae/localzen must stay direct), so add an explicit per-source proxy_url
that overrides http.ProxyFromEnvironment for just that source. Sources
without proxy_url keep the existing env/direct behavior.

Also accepts a User-Agent header per source (config already supported
headers) so CF-fronted resellers can be reached.

Verified e2e: tabitoken and justwoker now return tool_calls through
127.0.0.1:7890 (clash); trae stays direct. Full go test -tags luajit
passes.
2026-09-05 09:46:36 +08:00
4d197e4bd3 fix(trae): recover legacy [Called tool:...] tool call format
Root cause: trae-local-api is deployed to users without the "Fold past
assistant tool_calls into [Called tool: name({...})]" text format that
their own histories already contained. Trae-Local-API-LLM then mimics this
format in subsequent responses. The trae.lua parser only recognized
<tool_call>...</tool_call> or <toolcall>...</toolcall> tags, so
[Called tool: ...] responses were left unparsed and the client received
plain text where a structured tool_calls array should be.

Fix: Add a legacy pattern match at the END of parse_text_tool_calls to
catch the [Called tool: name({args})] shape and emit proper tool_calls.
This is a fallback; models should emit <tool_call> tags per system prompt,
but we tolerate the mimicked form for robustness.
2026-09-05 09:46:36 +08:00
3806aaee03 docs(workflow): codify the branch model — main / feature / release branches
Adopt GitHub Flow + release branches, replacing "everything straight to main
plus a tag" which caused the 1.4.2 pain (a fix had to be retro-fitted to the
released version, forcing a remote-tag delete + full re-upload).

- main: only long-lived branch, always deployable, accumulates the next version
- feature/<desc>: born from main, merged back when done
- release/vX.Y.Z: cut from main, tagged, installers built from the tag
- hotfixes land on the release branch AND are cherry-picked back to main so
  main never loses a fix
- end of lifecycle = retire the release branch (delete; or keep for long-term
  maintenance), no wholesale merge back — hotfixes already flowed
- explicitly no rebase of main, no release-branch-merge, no quick edits on main

Companion release checklist includes the upload lessons (PUT --http1.1) and the
replace-artifacts-by-deleting-the-tag catch.

Docs in docs/git-workflow.md (zh) and docs/git-workflow-en.md, linked from both
READMEs.
2026-08-31 12:36:17 +08:00
5b628b4d6f fix(scheduler): rebalance the pref score so metered sources aren't written off silently
Report: an allowance-metered source (sensenova) had its deepseek model sitting at
pref -10 while the WebUI showed zero failures. All three numbers were accurate,
and they exposed three compounding problems.

1. Quota exhaustion cost the SAME score as a real failure. RecordQuotaExhausted
   intentionally avoids failCount (an exhausted allowance is not a fault), so the
   UI showed fails=0 / cooling=false — yet it deducted the full prefFailStep (5).
   For a metered source, running out of budget is an everyday event, so the score
   drifted deep negative with no visible cause. Now quota and 429 events cost
   prefQuotaStep (1): the cooldown already keeps the slot out of rotation until
   the window resets, the score only needs a mild preference for slots with budget.

2. Recovery was 5:1 asymmetric. A failure cost -5 but a success only +1, so a
   slot at -10 needed ten consecutive successes just to reach neutral — which it
   could never get, because a low score makes the scheduler not pick it in the
   first place (starvation). Success now rewards prefSuccessStep (2): recovery
   from -10 needs five successes, while a real failure still outweighs one.

3. No idle decay. A penalised slot kept its negative score forever once it stopped
   being selected. Pref() now applies lazy decay: after prefDecayAfter (2 min) of
   no outcome, the score drifts one step back toward 0 per interval (never past
   0, never touches positive scores). Applied in Pref() and TryProbe(), so a
   naturally-recovered idle slot is schedulable again without needing a probe.

The failure penalty itself is unchanged (prefFailStep=5), so genuinely broken
upstreams are still marked as clearly worse than healthy ones.

Tests: quota penalty lighter than failure; 5 quota resets stay well above the
floor with failCount untouched; 429 is quota-class; recovery from -10 needs <=5;
idle decay rehabilitates a written-off slot, stops at 0, and never drags a
positive score; decay repeatedly lifts a slot off the prefMin floor. Updated the
two pre-existing tests that asserted the old -5/129 +1 values.
v1.4.2
2026-08-31 11:50:47 +08:00
2e3d5b79ad fix(adapters): stop dropping non-streaming tool calls (agent loops died on turn 2)
Four adapters handled tool_calls in transform_stream_chunk but lost them in
transform_response, so any NON-streaming tool-using conversation broke on its
second request: the client received finish_reason:"tool_calls" with no
tool_calls payload, replayed an assistant message whose function
name/arguments were empty, and the upstream rejected the next turn with

    400 invalid tool_call function, function/name/arguments cannot be empty

The production audit trail shows 46 such failures on sensenova alone.

- sensenova.lua: forward message.tool_calls, decoding the arguments JSON string
  into an object as the unified shape expects.
- gemini.lua: collect functionCall parts from candidates[].content.parts. Also
  correct finish_reason, since Gemini reports "STOP" even when it emitted a
  function call and clients keyed on it treat that as a finished answer.
- ollama.lua: the field was initialized to an empty table and never filled;
  fill it and likewise correct done_reason "stop" -> "tool_calls".

trae is a different failure with the same symptom: trae-local-api's OpenAI
endpoint (/v1/chat/completions, src/server.js:353) never reads the request's
`tools` array — only its Anthropic endpoint does — so the relayed model is never
told the tool schema and instead PRINTS a <tool_call>{...}</tool_call> block into
content, leaving message.tool_calls null and finish_reason "stop". An OpenAI
client sees an ordinary completion and its agent loop ends mid-conversation.
trae.lua now recovers the structured call from that text, strips the block from
user-visible content, and corrects finish_reason. Both tag spellings
(<tool_call>/<toolcall>, the latter is what the same codebase's Anthropic prompt
asks for) and all three argument key names (arguments/params/input) are accepted.
This is a defensive fallback: fixing the upstream shim to honour `tools` remains
the real fix, since the model still guesses parameter names.

Tests: TestNonStreamToolCallsPreserved covers all ten OpenAI-shaped adapters,
TestGeminiNonStreamToolCalls and TestOllamaNonStreamToolCalls cover their native
shapes, TestTraeTextToolCallRecovery covers both tag spellings, prose around the
block, and asserts a plain text answer never gains tool_calls.

Verified end-to-end against mock upstreams reproducing each shape: a full
two-round agent loop (tool call -> tool result -> final answer) now completes for
both the structured and the text-emitted variants.
2026-08-31 10:23:35 +08:00
2ebc01e03b fix(build): nsis.7z is not a required artifact — Setup exe is the formal Windows product v1.4.1 2026-08-30 13:49:34 +08:00
11776024e9 chore(build): gitignore cmd/gui/.cache (docker electron-builder cache dir) 2026-08-30 10:56:01 +08:00
bf1932c6c5 chore(version): 1.4.0 -> 1.4.1 2026-08-30 10:52:48 +08:00
22ddf5a411 chore(build): drop ELECTRON_BUILDER_CACHE so the wine toolchain cache stays in the volume, not the workspace 2026-08-30 10:50:01 +08:00
3698546d14 fix(build): run Windows NSIS packaging inside docker; add artifact size gate
The Windows installer has been broken since 1.3.0: electron-builder's NSIS step
needs wine to generate the uninstaller, but the host's wine was amd64-only (no
i386 runtime -> empty syswow64 -> `error c0000135`), so electron-builder silently
wrote a 264 KB installer shell with no payload. No check caught it and the broken
exe shipped. 1.4.0 reproduced the same failure this session.

Two fixes:

1. win-builder image gains node + wine32/wine64 (+ i386 arch). dist-win-docker.sh
   now runs BOTH the core cross-build and `npx electron-builder --win nsis`
   inside docker (USE_SYSTEM_WINE=true -> the image's wine). The wine prefix is
   initialized on first run in the shared cache volume, so syswow64/ntdll.dll
   exists - the exact thing the host lacked. The host needs no mingw/wine/node.

2. packaging/verify-dist.sh: a size-floor gate for GUI artifacts (exe >= 5 MB,
   deb/rpm/nsis.7z >= 10 MB). Wired into `make gui-dist` and `make gui-win-docker`
   and the dist-win-docker script, so a degenerate installer fails the build
   instead of reaching a Release. Verified: it rejects the 264 KB exe and passes
   the healthy artifacts.
2026-08-30 10:49:45 +08:00