mirror of
https://gitcode.com/JianFeeeee/ModelRouter.git
synced 2026-10-06 15:37:42 +00:00
refactor(quota): 配额改为按模型,删除整钥总配额
用户明确要求:配额应当是密钥对应的**每个模型的单独配额**,而非整体配额。
## 语义变更
删除 GWKey.TokenQuota / ReqQuota / Period / Hours(整钥总额)。
ModelScope 新增 ReqQuota —— 请求数配额下沉到每条模型范围。
现在:每条 models[] 各自带 token 配额 + 请求数配额 + 重置周期,
彼此独立。一个模型用满只影响该模型。
★ 为什么不保留整钥总额:它会让「把 A 模型的额度挪给 B」变成一次全局
重分配;按模型独立计费则每个模型各自可控,运维能直接看出哪个模型在吃预算。
## 连带改动
- checkQuota 合并 key 级与 scope 级判定;checkKeyQuotaRetry 整体删除
(顺带修掉上轮遗留的双重判定:入口不再先判空再重算)
- core:CreateKeyWithQuota / UpdateKeyWithQuota / ApplyQuota 全部删除,
改由 ValidateScopeQuotas 校验每条 scope 的配额
- admin key:scope 上的配额不强制(admin 的 scope 仍限制模型范围,
但不强制配额)—— 否则管理员会把自己锁在门外
- /api/v1/keys 不再回显 key 级配额字段(scope 里已含)
- WebUI:删除整钥配额徽标 / 「配额」按钮 / 创建表单的配额组 /
putScope 的整钥回传;模型砖块与范围编辑器新增「请求数配额」输入,
徽标显示 `1.0K 77×·1h`(未设配额显示 ∞)
## 判据
- TestOneModelsQuotaDoesNotBlockAnother 是本次核心保证。
★ 它第一版是**假判据**:m2 从不消耗,key-wide 计数器与 m1 自己的计数器
读数恰好相同,退回 key-wide 仍通过。变异测试抓到后改为「先用 m2 花掉
远超 m1 配额的量,再验证 m1 仍可用」—— 这样两种设计才可区分。
- TestUncappedModelNeverBlocked / TestAdminKeyScopesAreNotEnforced 新增
- UI 契约判据重写:整钥配额界面必须彻底消失(13 个符号)、
scope 编辑器必须往返 req_quota、putScope 只发 scope 列表
- 错误消息点名具体模型(TestKeyAPIRejectionNamesTheModel)
- 3/3 变异全被抓
实测(真实进程 + 浏览器):m2 配额 500000 连打 25 次全成功,
m1 配额 1000 立即 429「token quota exceeded for "m1" (4315/1000)」,
此后 m2/m3 仍 200。UI:整钥配额元素全为 0,砖块各显配额,
编辑器预填/保存正确,零 JS 异常。
(cherry picked from commit c51066f0b6)
This commit is contained in:
73
README_EN.md
73
README_EN.md
@ -38,10 +38,11 @@ Extracted and independently evolved from the multi-source LLM adapter layer of
|
||||
`reasoning_content`, `tool_calls`, `usage`).
|
||||
- **Image generation**: `POST /v1/images/generations`, routed to models with
|
||||
`kind: image`.
|
||||
- **Per-key usage quota**: each key carries its own token and request caps plus a
|
||||
reset period (hour/week/month/custom N hours), shared across every model that
|
||||
key may use. Exhaustion answers 429 + `Retry-After` so a client resumes when
|
||||
the window rolls over; admin keys are never capped.
|
||||
- **Per-model quota**: each key gives every model its own token and request caps
|
||||
plus a reset period (hour/week/month/custom N hours). One model running out
|
||||
affects only that model — the key's other models keep working. Exhaustion
|
||||
answers 429 + `Retry-After` naming the model, so a client resumes when the
|
||||
window rolls over; admin keys are never capped.
|
||||
- **Multimodal**: `content` arrays (`image_url` etc.) pass through losslessly;
|
||||
Anthropic/Gemini/Ollama are translated automatically.
|
||||
- **LuaJIT VM**: golua-binding LuaJIT; each adapter has its own VM + worker
|
||||
@ -175,56 +176,62 @@ under the `keys` field of the runtime file (encrypted at rest):
|
||||
delete the seed key.
|
||||
- The WebUI **Keys page** creates/deletes keys. Each key has a role (`admin`
|
||||
manages everything, `user` sees only its own key) and an optional **model
|
||||
scope** (model + source + token quota + reset period). Key cards show the
|
||||
key-wide caps as a badge (e.g. `250.0K·6h` / `77×·6h`); a **Quota** button
|
||||
edits the total token / request budget and its reset period, and the create
|
||||
form takes a budget too (those fields disable themselves for `admin`, which
|
||||
is never capped). The "My key" view shows a user its own budget.
|
||||
scope** (model + source + token quota + reset period). Each model brick shows
|
||||
its own budget badge (e.g. `1.0K 77×·1h`, `∞` when uncapped); clicking it
|
||||
edits that model's token quota, request quota and reset period. Quotas apply
|
||||
per model, so one model running out never blocks the key's others. The
|
||||
"My key" view shows a user its model scopes and their budgets.
|
||||
- Clients authenticate with any authorized key's plaintext as
|
||||
`Authorization: Bearer <key>`.
|
||||
- Deleting a key removes it from the store immediately.
|
||||
|
||||
#### Per-key usage quota
|
||||
|
||||
Each key can cap its own spend and reset period. Two levels apply at once:
|
||||
**Quotas are per model.** Each key's `models[]` list gives every model its own
|
||||
token budget and request budget. One model running out affects only that model
|
||||
— the key's other models keep working.
|
||||
|
||||
```yaml
|
||||
keys:
|
||||
- key: sk-gw-<hex>
|
||||
role: user
|
||||
name: agent-alice
|
||||
# ---- key-wide (across every model) ----
|
||||
token_quota: 5000000 # total token budget for this window, 0 = unlimited
|
||||
req_quota: 20000 # requests per window, 0 = unlimited
|
||||
period: nhour # "" | hour | week | month | nhour
|
||||
hours: 6 # n-hour only: resets every 6 hours
|
||||
# ---- per-model scope (optional) ----
|
||||
models:
|
||||
- model: m1
|
||||
token_quota: 1000000
|
||||
period: hour
|
||||
- model: deepseek-v4-flash
|
||||
token_quota: 1000000 # this key's token budget for this model
|
||||
req_quota: 20000 # requests within the window
|
||||
period: nhour # "" | hour | week | month | nhour
|
||||
hours: 6 # n-hour only
|
||||
- model: AUTO
|
||||
token_quota: 5000000 # AUTO is a quota entry like any other
|
||||
period: hour
|
||||
- model: kimi-k3 # no quota listed = unlimited
|
||||
```
|
||||
|
||||
- **There is deliberately no key-wide total.** A key-wide cap would make
|
||||
"move A's budget to B" a global reallocation; per-model budgets keep each
|
||||
model independently controllable, so it stays visible which model is
|
||||
actually consuming the spend.
|
||||
- `period`: empty = never resets (lifetime total); `hour` / `week` / `month` =
|
||||
fixed windows; `nhour` + `hours` = a custom hour count. **A misspelled
|
||||
period is rejected at write time** rather than silently becoming a
|
||||
never-resetting quota.
|
||||
- The key-wide cap is one budget shared by every model the key may use;
|
||||
quotas under `models[]` are counted per model. Both are isolated per key —
|
||||
one key's traffic never drains another's budget.
|
||||
- Quotas are isolated per key *and* per model within a key: one key exhausting
|
||||
`m1` never draws on another key's budget, and never blocks the same key's
|
||||
`m2`.
|
||||
- Usage counts chat, streaming and image requests, and survives a restart by
|
||||
replaying the audit log (40 days retained, covering the longest `month`
|
||||
window).
|
||||
- An exhausted quota returns **429 + `Retry-After`**
|
||||
(`rate_limit_exceeded`) so a client resumes when the window rolls over; a
|
||||
model the key may not use stays 403. **Admin keys are never capped**, so a
|
||||
cap can never lock the operator out.
|
||||
(`rate_limit_exceeded`) and the message names the model that ran out, so a
|
||||
client resumes when the window rolls over; a model the key may not use stays
|
||||
403. **Admin keys are never capped** (quotas on their scopes are not
|
||||
enforced either), so a cap can never lock the operator out.
|
||||
- Buckets are whole unix hours, so a window frees up at most an hour late
|
||||
(deliberately freeing late rather than overspending).
|
||||
- On `PUT /api/keys/{key}` the quota fields are optional: omitting them keeps
|
||||
the stored caps, sending `0` explicitly lifts a cap. Editing only the model
|
||||
scope never clears a budget that was already set.
|
||||
- `PUT /api/keys/{key}` submits quotas by submitting `models` — the caps are
|
||||
part of the scope, so there is no second budget that can drift out of sync
|
||||
with the model list. An explicit `0` lifts that model's cap.
|
||||
|
||||
##### Quota rejection vs capacity rejection
|
||||
|
||||
@ -247,14 +254,10 @@ back off concurrency or switch sources.
|
||||
Buckets are kept per (key, model, whole unix hour) for 40 days. Measured on an
|
||||
AMD 7840HS:
|
||||
|
||||
- Quota check per request: **149 ns** (caps set) / **42.6 ns** (no caps — it
|
||||
only looks up the key record and never touches a bucket) / **37 ns** (admin
|
||||
key returns immediately) — all **0 allocations**. Keys without caps cost
|
||||
almost nothing, so creating many of them is safe.
|
||||
- Recording one request: 283 ns, 3 allocations (unchanged from before this
|
||||
feature; the allocations come from the record ring buffer).
|
||||
- A window query scans the window, not the whole retention: 49 ns for 1 h,
|
||||
55 ns for 24 h, 3.9 µs for 30 d.
|
||||
- Recording one request: 283 ns, 3 allocations (unchanged from before this
|
||||
feature; the allocations come from the record ring buffer).
|
||||
- Memory: the production shape (7 keys x 8 models x 2 sources at full 40-day
|
||||
retention) costs about **3.7 MB**. The `source::model` bucket used by a
|
||||
source-pinned quota is created **lazily** — it is maintained only once some
|
||||
@ -267,7 +270,7 @@ AMD 7840HS:
|
||||
```
|
||||
HTTP/1.1 429 Too Many Requests
|
||||
Retry-After: 2100
|
||||
{"error":{"type":"rate_limit_exceeded","message":"key token quota exceeded (5000000/5000000, resets every 6h)"}}
|
||||
{"error":{"type":"rate_limit_exceeded","message":"token quota exceeded for \"deepseek-v4-flash\" (5000000/5000000)"}}
|
||||
```
|
||||
|
||||
### Model routing
|
||||
|
||||
Reference in New Issue
Block a user