mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-05 23:17:24 +00:00
feat(webui): 密钥配额表单 + 修复弹窗关闭错对象
功能:让 per-key 配额在 WebUI 里可配置可见,之前的实现只有 API 与
config.yaml 能配。
- 密钥卡片头部显示配额徽标(token / 请求数 + 重置窗口),admin key
不显示编辑入口(服务端本就永不受限,给入口只会让人以为配了会生效)。
- 新增「配额」编辑弹窗:token 配额、请求数配额、重置周期(复用既有
的 period 词表与 n-hour 联动),预填从 canvas 的 data-* 读。
- 创建密钥弹窗同步加配额字段;选 admin 角色时自动禁用(同样因为服务端
忽略 admin 的配额)。
- 「我的密钥」页新增 KEY-WIDE QUOTA 列,用户能看到自己这把 key 的预算。
修一个真 bug:保存弹窗用 $("#modal-wrap") 关闭自己,而全站弹窗共用这个
id、且可以叠加(seed key 提示就盖在密钥页上)。实测(共享 Chromium
CDP,seed 提示与配额弹窗共存)确认:保存后被移除的是 seed 提示,配额表单
反而留在屏幕上 —— 症状是「保存了但弹窗没关」,指向的方向完全错。改为用
点击的按钮 btn.closest("#modal-wrap") 解析自己的弹窗。createKey 有同样
问题,一并修。既有文件里另有 7 处同样写法,未动(不在本次范围,且新判据
只对本次改的两处断言,避免误伤)。
判据新增 internal/gateway/ui_quota_contract_test.go(6 例):
- 两个表单必须用 .closest 解析自己的弹窗
- putScope 必须带上 4 个配额字段(API 视其为指针,省略=清空预算)
- 创建请求必须真的发出配额字段
- **数据流判据**:徽标要真读 k.token_quota 等、编辑表单要真读
canvas 写的 data-kquota 等。只查字面量存在会漏 —— 字段躺在死分支里
判据照样通过(这是本轮实际踩到的:keyCapBadges 经 keyPeriodSuffix
间接读 k.period,被判据抓到后我把读取显式化而不是放宽判据)
- 弹窗扫描先剥注释,否则修复说明里引用的字面量会被当成违规
- 复用既有 ui_contract_test.go 的 jsFunctionBody(大括号配平);
自己第一版用 2000 字符固定窗口,被长注释顶开后仍在窗口外命中后面
函数的同名字段,读起来像通过 —— 窗口法在这里是假判据
7 个变异全部被抓(unsafe 关闭、putScope 丢字段、createKey 丢字段、
徽标不读字段、canvas 不写 data-*、kq-hours 改名、周期词表缺项)。
浏览器实测(共享 Chromium CDP,真实进程 + 加密配置):
- 徽标渲染 1.0K·1h / 5×·1h;编辑框预填 1000/5/hour,hours 框按周期联动
- 保存后回读 250000/77/nhour/6,徽标更新为 250.0K·6h,toast Saved
- 零 JS 异常
- **关键回归**:编辑模型 scope 后配额仍是 250000/77/nhour,未被清空
- 创建带配额的 key,服务端确认 {t:50000,r:300,p:week,role:user}
- user 视角「我的密钥」显示 777·1h 与 9×·1h
文档:README.md / README_EN.md 补「密钥用量配额」小节(配置示例、
周期词表、429 语义、admin 豁免、整点分桶最晚晚 1 小时释放、PUT 的
省略 vs 0 语义、429 响应样例),特性列表各加一条。
This commit is contained in:
52
README_EN.md
52
README_EN.md
@ -38,6 +38,10 @@ Extracted and independently evolved from the multi-source LLM adapter layer of
|
||||
`reasoning_content`, `tool_calls`, `usage`).
|
||||
- **Image generation**: `POST /v1/images/generations`, routed to models with
|
||||
`kind: image`.
|
||||
- **Per-key usage quota**: each key carries its own token and request caps plus a
|
||||
reset period (hour/week/month/custom N hours), shared across every model that
|
||||
key may use. Exhaustion answers 429 + `Retry-After` so a client resumes when
|
||||
the window rolls over; admin keys are never capped.
|
||||
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
|
||||
Anthropic/Gemini/Ollama are translated automatically.
|
||||
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
|
||||
@ -176,6 +180,54 @@ under the `keys` field of the runtime file (encrypted at rest):
|
||||
`Authorization: Bearer <key>`.
|
||||
- Deleting a key removes it from the store immediately.
|
||||
|
||||
#### Per-key usage quota
|
||||
|
||||
Each key can cap its own spend and reset period. Two levels apply at once:
|
||||
|
||||
```yaml
|
||||
keys:
|
||||
- key: sk-gw-<hex>
|
||||
role: user
|
||||
name: agent-alice
|
||||
# ---- key-wide (across every model) ----
|
||||
token_quota: 5000000 # total token budget for this window, 0 = unlimited
|
||||
req_quota: 20000 # requests per window, 0 = unlimited
|
||||
period: nhour # "" | hour | week | month | nhour
|
||||
hours: 6 # n-hour only: resets every 6 hours
|
||||
# ---- per-model scope (optional) ----
|
||||
models:
|
||||
- model: m1
|
||||
token_quota: 1000000
|
||||
period: hour
|
||||
- model: AUTO
|
||||
```
|
||||
|
||||
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
|
||||
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
|
||||
period is rejected at write time** rather than silently becoming a
|
||||
never-resetting quota.
|
||||
- The key-wide cap is one budget shared by every model the key may use;
|
||||
quotas under `models[]` are counted per model. Both are isolated per key —
|
||||
one key's traffic never drains another's budget.
|
||||
- Usage counts chat, streaming and image requests, and survives a restart by
|
||||
replaying the audit log (40 days retained, covering the longest `month`
|
||||
window).
|
||||
- An exhausted quota returns **429 + `Retry-After`**
|
||||
(`rate_limit_exceeded`) so a client resumes when the window rolls over; a
|
||||
model the key may not use stays 403. **Admin keys are never capped**, so a
|
||||
cap can never lock the operator out.
|
||||
- Buckets are whole unix hours, so a window frees up at most an hour late
|
||||
(deliberately freeing late rather than overspending).
|
||||
- On `PUT /api/keys/{key}` the quota fields are optional: omitting them keeps
|
||||
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
|
||||
scope never clears a budget that was already set.
|
||||
|
||||
```
|
||||
HTTP/1.1 429 Too Many Requests
|
||||
Retry-After: 2100
|
||||
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
|
||||
```
|
||||
|
||||
### Model routing
|
||||
|
||||
`/v1/chat/completions` `model` resolution order:
|
||||
|
||||
Reference in New Issue
Block a user