refactor(quota): 配额改为按模型,删除整钥总配额

用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。

## 语义变更

删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。

现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。

★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。

## 连带改动

- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
  (顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
  改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
  但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
  putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
  徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)

## 判据

- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
  ★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
  读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
  远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
  scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓

实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。

(cherry picked from commit c51066f0b6)
This commit is contained in:
JianFeeeee
2026-09-27 19:02:13 +08:00
parent cc5e725226
commit ebe60028f5
11 changed files with 515 additions and 645 deletions

View File

@ -38,10 +38,11 @@ Extracted and independently evolved from the multi-source LLM adapter layer of
`reasoning_content`, `tool_calls`, `usage`).
- **Image generation**: `POST /v1/images/generations`, routed to models with
`kind: image`.
- **Per-key usage quota**: each key carries its own token and request caps plus a
reset period (hour/week/month/custom N hours), shared across every model that
key may use. Exhaustion answers 429 + `Retry-After` so a client resumes when
the window rolls over; admin keys are never capped.
- **Per-model quota**: each key gives every model its own token and request caps
plus a reset period (hour/week/month/custom N hours). One model running out
affects only that model — the key's other models keep working. Exhaustion
answers 429 + `Retry-After` naming the model, so a client resumes when the
window rolls over; admin keys are never capped.
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
Anthropic/Gemini/Ollama are translated automatically.
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
@ -175,56 +176,62 @@ under the `keys` field of the runtime file (encrypted at rest):
delete the seed key.
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
manages everything, `user` sees only its own key) and an optional **model
scope** (model + source + token quota + reset period). Key cards show the
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
edits the total token / request budget and its reset period, and the create
form takes a budget too (those fields disable themselves for `admin`, which
is never capped). The "My key" view shows a user its own budget.
scope** (model + source + token quota + reset period). Each model brick shows
its own budget badge (e.g. `1.0K 77×·1h`, `∞` when uncapped); clicking it
edits that model's token quota, request quota and reset period. Quotas apply
per model, so one model running out never blocks the key's others. The
"My key" view shows a user its model scopes and their budgets.
- Clients authenticate with any authorized key's plaintext as
`Authorization: Bearer <key>`.
- Deleting a key removes it from the store immediately.
#### Per-key usage quota
Each key can cap its own spend and reset period. Two levels apply at once:
**Quotas are per model.** Each key's `models[]` list gives every model its own
token budget and request budget. One model running out affects only that model
— the key's other models keep working.
```yaml
keys:
- key: sk-gw-<hex>
role: user
name: agent-alice
# ---- key-wide (across every model) ----
token_quota: 5000000 # total token budget for this window, 0 = unlimited
req_quota: 20000 # requests per window, 0 = unlimited
period: nhour # "" | hour | week | month | nhour
hours: 6 # n-hour only: resets every 6 hours
# ---- per-model scope (optional) ----
models:
- model: m1
token_quota: 1000000
period: hour
- model: deepseek-v4-flash
token_quota: 1000000 # this key's token budget for this model
req_quota: 20000 # requests within the window
period: nhour # "" | hour | week | month | nhour
hours: 6 # n-hour only
- model: AUTO
token_quota: 5000000 # AUTO is a quota entry like any other
period: hour
- model: kimi-k3 # no quota listed = unlimited
```
- **There is deliberately no key-wide total.** A key-wide cap would make
"move A's budget to B" a global reallocation; per-model budgets keep each
model independently controllable, so it stays visible which model is
actually consuming the spend.
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
period is rejected at write time** rather than silently becoming a
never-resetting quota.
- The key-wide cap is one budget shared by every model the key may use;
quotas under `models[]` are counted per model. Both are isolated per key —
one key's traffic never drains another's budget.
- Quotas are isolated per key *and* per model within a key: one key exhausting
`m1` never draws on another key's budget, and never blocks the same key's
`m2`.
- Usage counts chat, streaming and image requests, and survives a restart by
replaying the audit log (40 days retained, covering the longest `month`
window).
- An exhausted quota returns **429 + `Retry-After`**
(`rate_limit_exceeded`) so a client resumes when the window rolls over; a
model the key may not use stays 403. **Admin keys are never capped**, so a
cap can never lock the operator out.
(`rate_limit_exceeded`) and the message names the model that ran out, so a
client resumes when the window rolls over; a model the key may not use stays
403. **Admin keys are never capped** (quotas on their scopes are not
enforced either), so a cap can never lock the operator out.
- Buckets are whole unix hours, so a window frees up at most an hour late
(deliberately freeing late rather than overspending).
- On `PUT /api/keys/{key}` the quota fields are optional: omitting them keeps
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
scope never clears a budget that was already set.
- `PUT /api/keys/{key}` submits quotas by submitting `models` — the caps are
part of the scope, so there is no second budget that can drift out of sync
with the model list. An explicit `0` lifts that model's cap.
##### Quota rejection vs capacity rejection
@ -247,14 +254,10 @@ back off concurrency or switch sources.
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
AMD 7840HS:
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
only looks up the key record and never touches a bucket) / **37 ns** (admin
key returns immediately) — all **0 allocations**. Keys without caps cost
almost nothing, so creating many of them is safe.
- Recording one request: 283 ns, 3 allocations (unchanged from before this
feature; the allocations come from the record ring buffer).
- A window query scans the window, not the whole retention: 49 ns for 1 h,
55 ns for 24 h, 3.9 µs for 30 d.
- Recording one request: 283 ns, 3 allocations (unchanged from before this
feature; the allocations come from the record ring buffer).
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
retention) costs about **3.7 MB**. The `source::model` bucket used by a
source-pinned quota is created **lazily** — it is maintained only once some
@ -267,7 +270,7 @@ AMD 7840HS:
```
HTTP/1.1 429 Too Many Requests
Retry-After: 2100
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
```
### Model routing